A system and method for embedding watermarks in chat content

By building symbol tables and splitter tables in chat content, analyzing the embedding locations and using symbol recognition modules, the problem of watermarks being easily removed and untraceable is solved, and the watermarks being embedded and traced from the source without affecting the display.

CN113987430BActive Publication Date: 2025-07-22JIANGSU XINHE YIJIA INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111306603.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-07-22
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

When the prior art embeds watermarks in chat content, it is easy to be removed by the user, affecting the viewing experience, and the source cannot be traced after copying.

Method used

By building symbol tables and splitter tables, analyzing chat content to find suitable embedding locations, embed user IDs using symbols and trace the source through symbol identification modules, ensuring that the watermark does not affect the content display.

Benefits of technology

It realizes that the watermark is embedded in the chat content without affecting the display, and the information source can still be traced after the user takes a screenshot or copyes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113987430B_ABST
    Figure CN113987430B_ABST
Patent Text Reader

Abstract

The present invention relates to a system for embedding watermarks in chat content, including a symbol table and delimiter table module for screening special symbols, finding symbols that have no special meaning and do not affect users' reading of chat content, and constructing a symbol table and a delimiter table; a content analysis module for analyzing chat content to find suitable embedding positions where embedding symbols will not affect users' reading of chat content; a symbol embedding module for converting a user ID into a symbol and embedding it in chat content according to certain rules; and a symbol recognition module for obtaining chat content with embedded symbols. The system and method for embedding watermarks in chat content analyze chat content, embed source information, complete watermark embedding while not affecting content display, add source information to chat content, prevent the problem that the fixed position of the watermark is easy to eliminate, and can also trace the source when information is leaked through screenshots, taking pictures or copying information by users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and specifically provides a system and method for embedding watermarks in chat content. Background Art

[0002] A watermark refers to a semi-transparent logo or icon added to a picture to prevent others from stealing the picture.

[0003] Existing technologies add pictures and text as watermarks to the background of chats. Although they control screenshots and photographed pictures and enable traceability of source information, users can use picture compilation software to remove the background watermark. At the same time, the watermarking technology may affect the viewing experience. When users copy chat content and use it elsewhere, the source of this content cannot be traced. Therefore, a system and method for embedding watermarks in chat content are proposed to solve the above problems. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides a system and method for embedding watermarks in chat content, which have the advantages of completing watermark embedding without affecting content display, and solve the problems that the watermarking technology may affect the viewing experience and the source of the content cannot be traced when users copy chat content and use it elsewhere.

[0005] The present invention provides the following technical solution: A system for embedding watermarks in chat content, comprising:

[0006] A symbol table and delimiter table module, configured to screen special symbols, find symbols that have no special meaning and do not affect users' reading of chat content, and construct a symbol table and a delimiter table;

[0007] A content analysis module, configured to analyze chat content, find suitable embedding positions, and not affect users' reading of chat content after embedding symbols;

[0008] A symbol embedding module, configured to convert a user ID into a symbol and embed it in chat content according to a certain rule;

[0009] A symbol recognition module, configured to obtain chat content with embedded symbols, extract the embedded symbols according to the set rule, and convert them into a user ID to achieve source traceability;

[0010] A server, configured to transmit and save chat content;

[0011] A user terminal, configured to receive chat content sent by the server and send chat content received by the server.

[0012] Another technical problem to be solved by the present invention is to provide a method for embedding watermarks in chat content, including the following steps:

[0013] 1) List the symbols suitable for embedding in chat content and construct a symbol table;

[0014] 2) List the symbols that are not commonly used and construct a delimiter table;

[0015] 3) Analyze the chat content to find the embedding positions;

[0016] 4) Obtain the user ID of the current chat user;

[0017] 5) Translate the binary into symbols and embed them in the chat content;

[0018] 6) After the user takes a screenshot or copies the text and shares it, the user can forward and share the content by taking a screenshot with the mobile phone or copying the text;

[0019] 7) The recognition module obtains the embedded symbols in the text;

[0020] 8) Obtain the user ID and determine the information source.

[0021] Further, the detailed content in step 1) is to screen special symbols, find symbols that have no meaning and do not affect the user's reading of the chat content as much as possible, and construct a symbol table. The detailed content in step 2) is to screen special symbols, find symbols that have no meaning and are not commonly used, and construct a delimiter table.

[0022] Further, the detailed content in step 3) is to perform semantic recognition to find fixed phrases, juxtaposed participles, auxiliary words, semantic words, punctuation marks, etc. in the chat content and calculate the number of occurrences of each.

[0023] Further, the detailed content in step 4) is as follows:

[0024] a. Through a network request, obtain the unique ID of the user;

[0025] b. According to the checksum calculation formula, calculate the checksum of this ID and append the checksum to the last digit of the user ID;

[0026] c. Convert the user ID into binary. The conversion technique is the decimal-to-binary method.

[0027] Further, the calculation method of the checksum is as follows:

[0028] First, multiply each digit of the user ID by itself respectively, then add the results of these multiplications, and then divide by 10 to see the remainder. The remainder can only be 0 - 1 - 2 - 3 - 4 - 5 - 6 - 7 - 8 - 9.

[0029] Furthermore, the embedding rules in step 5) are as follows:

[0030] a. Calculate the occurrence times of juxtaposed participles, auxiliary words, semantic words, punctuation marks, the beginning of a sentence, the end of a sentence, etc. through semantic recognition technology;

[0031] If the occurrence times of various types of words are greater than the binary length of the user, they can be selected as the embedding positions. If multiple positions meet the condition simultaneously, the position without symbols is preferred first to prevent interference with each other. If the embedding position already has the same symbol, this symbol will be removed;

[0032] If none of them meet the condition, calculate the total occurrence times of various types of words in the chat content. When it is greater than the binary length of the user, then select various types of words as the embedding positions. If the embedding position already has the same symbol, this symbol will be removed;

[0033] If it still does not meet the condition, calculate the length of the text content. When the length is greater than the binary length of the user, then select the positions between each character as the embedding positions. At this time, only space characters are suitable for embedding. If the embedding position already has the same symbol, this symbol will be removed;

[0034] Otherwise, no embedding is performed for this content;

[0035] b. Select the symbol to be embedded. Select the symbol to be embedded according to the position to be embedded in the content. If there are the same symbols at the position to be embedded, the same symbols will be deleted before embedding;

[0036] c. Embed the binary code. The position to be embedded needs to be greater than the binary length. 0 in the binary means no embedding is required, and 1 means embedding is required;

[0037] d. At the starting and ending positions of the embedded symbol, embed the starting symbol and the ending symbol into the content;

[0038] Determine the starting symbol (separator);

[0039] Determine the ending symbol (separator).

[0040] Furthermore, the detailed content in step 7) is as follows:

[0041] a. Convert the content in the screenshot into text through image recognition technology;

[0042] b. Determine the embedding position according to the separator;

[0043] c. Determine the symbol to be embedded according to the embedding position;

[0044] d. Identify juxtaposed participles, auxiliary words, semantic words, punctuation marks, etc. in the content;

[0045] e. According to the result of semantic recognition, read the binary ID of the current watermark.

[0046] Further, the detailed content of step 8) is as follows:

[0047] a. Convert the obtained QR code into a decimal value;

[0048] b. Use the last digit of the decimal value as the check digit, and use the remaining digits of the decimal as the user ID;

[0049] c. Use the check code algorithm to calculate the check code value of the user ID;

[0050] d. If the check code value is different from the check digit, this watermark is modified and discarded. If the check code value is the same as the check digit, this watermark is valid;

[0051] e. Determine the information source according to the user ID and trace back to the content source.

[0052] Compared with the prior art, the technical solution of the present application has the following beneficial effects:

[0053] The system and method for embedding a watermark in chat content analyze the chat content, embed the source information, complete the watermark embedding without affecting the content display, add the source information to the chat content, prevent the problem that the watermark position is fixed and easy to be eliminated, and can also trace the source when the user leaks information by taking screenshots, taking pictures or copying information. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the system of the present invention;

[0055] Figure 2 It is a schematic diagram of information transmission of the present invention;

[0056] Figure 3 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0057] Hereinafter, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Please refer to Figure 1 , the system for embedding a watermark in chat content includes:

[0059] Symbol table and delimiter table module, which is used to screen special symbols, find symbols that have no special meaning and do not affect the user's reading of the chat content, and construct the symbol table and the delimiter table;

[0060] Content analysis module, which is used to analyze the chat content, find suitable embedding positions, and will not affect the user's reading of the chat content after embedding symbols;

[0061] Symbol embedding module, which is used to convert the user ID into symbols and embed them in the chat content according to certain rules;

[0062] Symbol recognition module, which is used to obtain the chat content with embedded symbols, extract the embedded symbols according to the set rules, and convert them into user IDs to achieve source tracing;

[0063] Server, which is used to transmit and save the chat content;

[0064] User side, which is used to receive the chat content sent by the server and to send the chat content received by the server.

[0065] Please refer to Figures 2-3 , the original text of the chat content is transmitted and saved on the server side, and only when it is displayed on the front end, the watermark of the user ID is added. Specifically, it includes the following steps:

[0066] 1) List the symbols suitable for embedding in the chat content and construct the symbol table;

[0067] 2) List the symbols that are not commonly used and construct the delimiter table;

[0068] 3) Analyze the chat content and find the embedding positions;

[0069] 4) Obtain the user ID of the current chat user;

[0070] 5) Translate the binary into symbols and embed them in the chat content;

[0071] 6) When the user takes a screenshot or copies the text and shares it, the user can forward and share the content by taking a screenshot with the mobile phone or copying the text;

[0072] 7) The recognition module obtains the embedded symbols in the text;

[0073] 8) Obtain the user ID and determine the information source.

[0074] Among them, the detailed content in step 1) is to screen special symbols, find symbols that have no meaning and do not affect the user's reading of the chat content as much as possible, and construct the symbol table.

[0075] For example:

[0076] Spaces, (.) in English, ~, -, etc.

[0077] Through semantic recognition technology, a large number of documents are processed for model training to find the symbols that are most frequently used in various scenarios, and a symbol table is constructed based on this. For example:

[0078] End of sentence Word segmentation Parallelism All positions 。 、 ; Space character

[0079] The detailed content in step 2) is to screen special symbols, find symbols that have no meaning and are not usually used, and construct a delimiter table.

[0080] For example:

[0081] Symbols such as ^, |, `, etc. are not usually used and can be used as delimiters.

[0082] The detailed content in step 3) is to perform semantic recognition to find fixed phrases, coordinate participles, auxiliary words, semantic words, punctuation marks, etc. in the chat content and calculate the number of times each appears.

[0083] It should be noted that semantic recognition and semantic analysis refer to using various methods to learn and understand the semantic content expressed by a text. Any understanding of language can be classified as the category of semantic analysis. The goal of semantic analysis is to achieve automatic semantic analysis of various language units (including words, sentences, and texts, etc.) by establishing effective models and systems, so as to understand the true semantics expressed by the entire text. Existing public technologies, such as: Baidu's NLP semantic analysis technology.

[0084] The detailed content in step 4) is as follows:

[0085] a. Through a network request, obtain the unique ID of the user;

[0086] b. According to the checksum calculation formula, calculate the checksum of this ID and append the checksum to the last digit of the user ID;

[0087] c. Convert the user ID into binary, and the conversion technology is the decimal-to-binary method.

[0088] The calculation method of the checksum is as follows:

[0089] First, multiply each digit of the user ID by itself, then add the results of these multiplications, and then divide by 10 to look at the remainder. The remainder can only be 0 - 1 - 2 - 3 - 4 - 5 - 6 - 7 - 8 - 9.

[0090] The embedding rules in step 5) are exemplified as follows:

[0091] a. Through semantic recognition technology, calculate the number of times coordinate participles, auxiliary words, semantic words, punctuation marks, the beginning of a sentence, the end of a sentence, etc. appear;

[0092] If the number of occurrences of various types of words is greater than the binary length of the user, it can be selected as the embedding position. If multiple positions meet the condition simultaneously, the position without symbols is preferred first to prevent interference with each other. If the embedding position already has the same symbol, this symbol will be removed;

[0093] When none of the above conditions are met, calculate the total number of occurrences of various types of words in the chat content. When it is greater than the binary length of the user, select various types of words as the embedding positions. If the embedding positions already have the same symbol, this symbol will be removed;

[0094] If it still does not meet the condition, calculate the length of the text content. When the length is greater than the binary length of the user, select the positions between each character as the embedding positions. At this time, only space characters are suitable for embedding. If the embedding positions already have the same symbol, this symbol will be removed;

[0095] Otherwise, no embedding is performed for this content;

[0096] Based on this, the start and end positions of the embedding can be determined. At the same time, according to the embedding positions, the delimiters are confirmed. The delimiters are usually uncommon symbols;

[0097] For example:

[0098] For juxtaposed participles, the delimiter is |;

[0099] At the end of a sentence, the delimiter is ^;

[0100] For all positions, the delimiter is `;

[0101] b. Select the symbol to be embedded. According to the positions to be embedded in the content, select the symbol to be embedded. If there are the same symbols at the positions to be embedded, the same symbols will be deleted before embedding;

[0102] c. Embed the binary code. The positions to be embedded need to be greater than the length of the binary. 0 in the binary means no embedding is required, and 1 means embedding is required;

[0103] d. At the start and end positions of the embedded symbol, embed the start symbol and the end symbol into the content;

[0104] Determine the start symbol (delimiter);

[0105] Determine the end symbol (delimiter);

[0106] For example:

[0107] The positions to be embedded are juxtaposed participles, and the delimiter is |;

[0108] The symbol to be embedded is 、;

[0109] Then determine the start symbol (|) and the end symbol (|).

[0110] The following are data examples:

[0111] ID Insertion position Separator Insertion symbol 1 Parallel participles | 、 2 End of sentence ^ 。 3 All positions ` Space character

[0112] The details in step 7) are as follows:

[0113] a. Convert the content in the screenshot into text through image recognition technology;

[0114] b. Determine the embedding position according to the delimiter;

[0115] c. Determine the embedded symbol according to the embedding position;

[0116] d. Identify the coordinating participles, auxiliary words, semantic words, punctuation marks, etc. in the content;

[0117] e. According to the result of semantic recognition, read the binary ID of this watermark.

[0118] The details of step 8) are as follows:

[0119] a. Convert the obtained QR code into a decimal value;

[0120] b. Use the last digit of the decimal value as the check digit, and the remaining digits of the decimal as the user ID;

[0121] c. Use the check code algorithm to calculate the check code value of the user ID;

[0122] d. If the check code value is different from the check digit, this watermark is modified and discarded. If the check code value is the same as the check digit, this watermark is valid;

[0123] e. Determine the information source according to the user ID and trace back to the content source.

[0124] Beneficial effects:

[0125] By analyzing the chat content and embedding the source information, the watermark embedding is completed without affecting the content display. Adding the source information to the chat content prevents the problem that the watermark position is fixed and easy to eliminate. When the user leaks information by taking screenshots, taking photos or copying information, the source can also be traced.

[0126] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0127] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, and the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for embedding a watermark in chat content, characterized in that, Specifically, it includes the following steps: 1) List the symbols suitable for embedding in the chat content and construct a symbol table; 2) List the infrequently used symbols and construct a delimiter table; 3) Analyze the chat content to find the embedding positions; 4) Obtain the user ID of the current chat user; 5) Convert the user ID into binary, translate the binary into symbols, and embed them in the chat content; 6) The user takes a screenshot or copies the text and then shares it; 7) The recognition module obtains the embedded symbols in the text; 8) Obtain the user ID and determine the information source; The embedding rule in step 5) is as follows: a. Calculate the number of occurrences of coordinate participles, auxiliary words, semantic words, punctuation marks, the beginning of a sentence, and the end of a sentence in the chat content group through semantic recognition technology; If the number of occurrences of each type of word is greater than the binary length of the user, select it as the embedding position. If multiple positions meet the condition simultaneously, select the position without a symbol. If there is already the same symbol at the embedding position, remove this symbol; If none of them meet the condition, calculate the total number of occurrences of each type of word in the chat content. When it is greater than the binary length of the user, select each type of word as the embedding position. If there is already the same symbol at the embedding position, remove this symbol; If it still does not meet the condition, calculate the length of the text content. When the length is greater than the binary length of the user, select the positions between each character as the embedding positions. At this time, only space characters are suitable for embedding. If there is already the same symbol at the embedding position, remove this symbol; Otherwise, no embedding is performed for this content; b. Select the symbol to be embedded. Select the symbol to be embedded according to the position to be embedded in the content. If there is already the same symbol at the position to be embedded, delete the same symbol before embedding; c. Embed the binary code. The position to be embedded needs to be greater than the length of the binary. 0 in the binary means no embedding is required, and 1 means embedding is required; d. At the start and end positions of the embedded symbol, embed the start symbol and the end symbol in the content; Determine the start symbol; Determine the end symbol; The start symbol and the end symbol are delimiters determined according to the embedding position.

2. The method for embedding a watermark in chat content according to claim 1, wherein: The detailed content in step 1) is to screen special symbols, find symbols that have no meaning and do not affect the user's reading of the chat content, and construct a symbol table. The detailed content in step 2) is to screen special symbols, find symbols that have no meaning and are usually not used, and construct a delimiter table.

3. A method for embedding a watermark in chat content according to claim 1, characterized in that: The detailed content in step 4) is as follows: a. Through a network request, obtain the unique ID of the user; b. According to the checksum calculation formula, calculate the checksum of this ID and append the checksum to the last digit of the user ID; c. Convert the user ID into binary. The conversion technology is the decimal-to-binary method.

4. A method for embedding a watermark in chat content according to claim 3, characterized in that: The calculation method of the checksum is as follows: First, multiply each digit of the user ID by itself respectively, then add the results of these multiplications. Next, divide by 10 and look at the remainder. The remainders are 0, 1, 2, 3, 4, 5, 6, 7, 8, 9.

5. A method for embedding a watermark in chat content according to claim 1, characterized in that: The detailed content in step 7) is as follows: a. Convert the content in the screenshot into text through image recognition technology; b. Determine the embedding position according to the delimiter; c. Determine the embedded symbol according to the embedding position; d. Identify coordinating participles, auxiliary words, semantic words, and punctuation marks in the content; e. Read the binary ID of the current watermark according to the result of semantic recognition.

6. A method for embedding a watermark in chat content according to claim 1, characterized in that: The detailed content of step 8) is as follows: a. Convert the obtained binary to a decimal value; b. Use the last digit of the decimal value as the check digit, and use the remaining digits of the decimal as the user ID; c. Use the checksum algorithm to calculate the checksum value of the user ID; d. If the checksum value is different from the check digit, the watermark is modified and discarded. If the checksum value is the same as the check digit, the watermark is valid; e. Determine the information source according to the user ID and trace back to the content source.

7. A system for implementing a method of embedding a watermark in chat content according to any one of claims 1-6, characterized in that, It includes: A symbol table and delimiter table module, which is used to screen special symbols, find symbols that have no special meaning and do not affect the user's reading of the chat content, construct a symbol table, find symbols that have no meaning and are not usually used, and construct a delimiter table; A content analysis module, which is used to analyze the chat content, find a suitable embedding position, and does not affect the user's reading of the chat content after embedding symbols; A symbol embedding module, which is used to convert the user ID into symbols and embed them in the chat content according to certain rules; A symbol recognition module, which is used to obtain the chat content with embedded symbols, extract the embedded symbols according to the set rules, convert them into user IDs, and achieve source tracing; A server, which is used to transmit and save the chat content; A user terminal, which is used to receive the chat content sent by the server and to send the chat content received by the server.

Citation Information

Patent Citations

  • Unicode coding-based text watermark embedding method and extraction method

    CN106570356A