Content generation method, generated content detection method and apparatus

By constructing character correspondences based on similar glyphs and watermark bitstreams, the issues of accuracy and robustness in content generation method recognition are resolved, enabling accurate judgment and isolation of generative content and preventing knowledge base contamination.

CN116595964BActive Publication Date: 2026-07-21ALIBABA (CHINA) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA (CHINA) CO LTD
Filing Date
2023-04-25
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of content generation method recognition is limited by the amount of content, has poor robustness, and cannot effectively isolate model-generated content from human-created content, leading to knowledge base pollution.

Method used

By constructing character correspondences that are the same or similar in glyphs across multiple languages, a content generation algorithm is used to generate target language content. Some or all of the target language characters are then converted into characters from other languages ​​with the same or similar glyphs, and a watermark bitstream is added to identify the content generation method.

Benefits of technology

It improves the recognition accuracy of generative content, enhances robustness to editing operations, and effectively isolates model-generated content from human-created content, preventing knowledge base contamination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116595964B_ABST
    Figure CN116595964B_ABST
Patent Text Reader

Abstract

The application discloses a content generation method and a generative content detection method. The content generation method comprises the following steps: constructing a character corresponding relationship of multiple languages with the same or similar characters; generating content in a target language through a content generation algorithm; and converting part or all of the target language characters in the content into other language characters with the same or similar characters according to the character corresponding relationship and content source information. In this way, the generative content is added with a generative content identification watermark through the conversion of homograph characters. Since the watermark is at the character level, the coverage of the watermark on the letter characters is large, so that the content generation method can be effectively identified with only a small amount of content, and even if the generative content is edited subsequently, the content generation method can still be effectively detected, and the algorithm-generated content can be effectively isolated from the human-created content, thereby avoiding pollution of the artificial knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to content generation methods and apparatus, generative content detection methods and apparatus, automatic question answering methods and apparatus, question answering methods and apparatus, creative idea generation methods and apparatus, program creation methods and apparatus, and electronic devices. Background Technology

[0002] Text content generation models have brought great convenience to human work and life. In some situations, it is necessary to determine whether a piece of content is automatically generated by a model or created by humans. Currently, the main approach is to extract various features of the content (such as writing style features) to determine whether the content comes from a certain type of model. However, this method has at least the following problems: 1) Detection based on content elements such as writing style features requires large sections of text for effective recognition; when the amount of content is small, the recognition accuracy is not high; 2) It has poor robustness; after the target content has undergone editing operations such as insertion, deletion, or modification, it is difficult to provide effective detection results; 3) It lacks content isolation capabilities, which can lead to the model-generated content polluting the human knowledge base. Summary of the Invention

[0003] This application provides a content generation method to address the problems in existing technologies where the accuracy of content generation method recognition is limited by the amount of content and the robustness of content generation method recognition for edited generated content is poor. This application also provides a content generation apparatus, a generated content detection method and apparatus, and an electronic device.

[0004] This application provides a content generation method, including:

[0005] Construct correspondences between characters that are identical or similar in glyphs across multiple languages;

[0006] Generate content in the target language using content generation algorithms;

[0007] Based on the character correspondence and the source information of the content, some or all of the target language characters in the content are converted into other language characters with the same or similar glyphs.

[0008] Optionally, the source information of the content is that the content is generated by an algorithm;

[0009] The step of converting some or all of the target language characters in the content into other language characters with the same or similar glyphs based on the character correspondence and the source information of the content includes:

[0010] Based on the character correspondence and the source information of the content, some or all of the target language characters are converted into any other language character with the same or similar glyphs.

[0011] Optionally, the source information of the content is intellectual property information; the method further includes:

[0012] Based on the character correspondence, construct the correspondence between homographs and binary numbers corresponding to characters in the target language;

[0013] The intellectual property information is converted into a watermarked bitstream;

[0014] The step of converting some or all of the target language characters in the content into other language characters with the same or similar glyphs based on the character correspondence and the source information of the content includes:

[0015] If the target language character has homographs, then the number of binary digits corresponding to the target language character is obtained based on the number of homographs of the target language character;

[0016] Based on the number of bits in the binary number, read the watermark bit corresponding to the number of bits from the watermark bit stream;

[0017] Based on the correspondence between the homograph and the binary number, obtain the target homograph corresponding to the watermark bit of the corresponding number of bits;

[0018] If the target homograph is a character from another language that has the same or similar glyph as the target language character, then the target language character is replaced with the target homograph.

[0019] Optionally, constructing the correspondence between homographs and binary numbers corresponding to characters in the target language based on the character correspondence includes:

[0020] Based on the character correspondence, obtain the number of homographs of the target language characters;

[0021] The number of bits in the binary number is determined based on the number of homographs.

[0022] Based on the number of bits in the binary number, construct the correspondence between the homograph and the binary number.

[0023] Optional, also includes:

[0024] The content converted from homographs corresponding to the content of the target language is stored in a knowledge base, which then performs content retrieval based on the character encoding corresponding to the content retrieval conditions.

[0025] This application also provides a generative content detection method, including:

[0026] Obtain content in the target language;

[0027] To obtain the correspondence between characters that are the same or similar in glyphs in multiple languages;

[0028] Based on the character correspondence, other language characters in the content of the target language are converted into target language characters with the same or similar glyphs;

[0029] If the content fragment composed of multiple target language characters after conversion has semantic information, then the content of the target language is determined to be generative content.

[0030] Optionally, if the content fragment composed of multiple target language characters after conversion has semantic information, then determining that the content of the target language is generative content includes:

[0031] Get the number of characters converted from the same character;

[0032] If the content fragment composed of multiple target language characters after conversion has semantic information and the number of characters is greater than the first threshold, then the content of the target language is determined to be generative content.

[0033] Optional, also includes:

[0034] Obtain the correspondence between homographs and binary numbers of characters in the target language;

[0035] Based on the correspondence between the homographs and binary numbers, obtain the binary numbers corresponding to the homographs in the content of the target language;

[0036] Add the binary number corresponding to the homograph to the watermark bitstream;

[0037] After the watermark information is extracted, the watermark bitstream is converted into intellectual property information.

[0038] This application also provides a generative content detection method, including:

[0039] To obtain the correspondence between characters that are the same or similar in glyphs in multiple languages;

[0040] Obtain content in the target language;

[0041] Based on the character correspondence, other language characters in the content of the target language are converted into target language characters with the same or similar glyphs;

[0042] The characters in the content that undergo homograph conversion are set as the corresponding first flag, the characters in the content that do not undergo homograph conversion are set as the corresponding second flag, and the non-homograph characters in the content are set as the third flag. Based on the position of the character in the content and the flag corresponding to the character, the character flag string corresponding to the content is obtained.

[0043] For the second flag segment in the character flag string, if the length of the second flag segment is less than the second threshold and the character segment corresponding to the flag segment adjacent to the second flag segment has semantic information, then the second flag segment and the character segment corresponding to the adjacent flag segment are used as generative content.

[0044] This application also provides a question-and-answer method, including:

[0045] Construct correspondences between characters that are identical or similar in glyphs across multiple languages;

[0046] Receive dialogue requests sent by the client;

[0047] Based on the dialogue information carried in the request, dialogue content is generated using a content generation algorithm;

[0048] Based on the character correspondence, some or all of the target language characters in the dialogue content are converted into other language characters with the same or similar glyphs;

[0049] The dialogue content after homograph conversion is sent back to the client.

[0050] This application also provides a method for answering questions, including:

[0051] Construct correspondences between characters that are identical or similar in glyphs across multiple languages;

[0052] Receive requests from clients to retrieve answers to specific questions;

[0053] The answer content corresponding to the target question is generated using a content generation algorithm;

[0054] Based on the character correspondence, some or all of the target language characters in the answer content are converted into other language characters with the same or similar glyphs;

[0055] The answer content after homograph conversion is sent back to the client.

[0056] This application also provides a method for generating creative ideas, including:

[0057] Construct correspondences between characters that are identical or similar in glyphs across multiple languages;

[0058] Receive requests from clients to retrieve creative content related to a specific creative question;

[0059] Creative content corresponding to the target creative problem is generated using a content generation algorithm;

[0060] Based on the character correspondence, some or all of the target language characters in the creative content are converted into other language characters with the same or similar glyphs;

[0061] The creative content, converted from homographs, is sent back to the client.

[0062] This application also provides a method for creating a program, including:

[0063] Construct correspondences between characters that are identical or similar in glyphs across multiple languages;

[0064] Receive programming requests sent by the client;

[0065] Based on the programming requirements information carried in the request, program code is generated using a content generation algorithm;

[0066] Based on the character correspondence, some or all of the target language characters in the program code are converted into other language characters with the same or similar glyphs;

[0067] The program code after converting homographs is sent back to the client.

[0068] This application also provides an electronic device, including:

[0069] Processor; and

[0070] A memory for storing a program that implements the method according to any one of the preceding claims, wherein the device is powered on and the program runs the method via the processor.

[0071] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the various methods described above.

[0072] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the various methods described above.

[0073] Compared with the prior art, this application has the following advantages:

[0074] The content generation method provided in this application constructs a character correspondence relationship between multiple languages ​​that is identical or similar in glyphs; generates content in the target language using a content generation algorithm; and converts some or all of the target language characters in the content into characters of other languages ​​with the same or similar glyphs based on the character correspondence relationship and the source information of the content. This method of adding a generative content watermark to the generative content through homograph conversion effectively increases the watermark capacity in the generative content because the watermark is at the character level and has a large coverage of alphabetic characters. Furthermore, the watermarked content, composed of character encodings from multiple languages, is significantly different from human-created content. Under this premise, even if the text volume of the generative content is small, it is possible to accurately determine whether the content has been watermarked to identify the content generation method, i.e., whether it was generated by the content generation algorithm. This allows for effective identification of the content generation method with only a small amount of content, preparing for a significant improvement in the accuracy of content generation method identification when the content volume is small. Furthermore, because the determination of watermark presence can be accurate down to the word and phrase level, it can effectively detect the generation method even after subsequent editing operations such as insertion, deletion, and modification of the generative content. It exhibits high robustness to editing operations not present in the original generative content. Moreover, by using homograph modulation, the inherent encoding and computer semantics of the generative content are disrupted while maintaining the human visual effect. This effectively isolates model-generated content from human-created content, preventing contamination of the human knowledge base.

[0075] The content generation method provided in this application involves: obtaining character correspondences in multiple languages ​​that are identical or similar in glyphs; obtaining content in a target language; converting characters from other languages ​​within the target language content into target language characters with identical or similar glyphs based on the character correspondences; and determining that the target language content is generative content if the content fragment composed of the converted target language characters has semantic information. This processing method allows for the rapid determination of the content generation method by converting homographs into target language characters and identifying the generation method of the target language content based on whether multiple target language characters have contextual semantics.

[0076] The content generation method provided in this application involves: acquiring character correspondences in multiple languages ​​that are identical or similar in glyphs; acquiring content in a target language; converting characters from other languages ​​in the target language content into target language characters with identical or similar glyphs based on the character correspondences; setting characters in the content that undergo homograph conversion as corresponding first flags, characters in the content that do not undergo homograph conversion as corresponding second flags, and non-homograph characters in the content as third flags; acquiring a character flag string corresponding to the content based on the position of the character in the content and the flags corresponding to the character; and for the second flag segment in the character flag string, if the length of the second flag segment is less than a second threshold and the character segment corresponding to the flag segment adjacent to the second flag segment has semantic information, then the second flag segment and the character segment corresponding to the adjacent flag segment are used as generative content. This processing method allows for the identification of which segment of information in the target language content is automatically generated by the algorithm and which segment is created by humans, thus effectively improving the accuracy of content generation method recognition. Attached Figure Description

[0077] Figure 1 A flowchart illustrating an embodiment of the content generation method provided in this application;

[0078] Figure 2 A schematic diagram illustrating a scenario of an embodiment of the content generation method provided in this application;

[0079] Figure 3 A flowchart illustrating an embodiment of the generative content detection method provided in this application. Detailed Implementation

[0080] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0081] This application provides a content generation method and apparatus, a generative content detection method and apparatus, and an electronic device. The various solutions are described in detail below in each embodiment.

[0082] First Embodiment

[0083] Please refer to Figure 1 This is a flowchart of the content generation method of this application. In this embodiment, the method may include the following steps:

[0084] Step S101: Obtain the corresponding relationships of characters that are the same or similar in glyphs in multiple languages.

[0085] Artificial characters are stored and transmitted by computers in one or more sets of unified encoding methods. Each text character is equivalent to a fixed encoding number for a computer. For example, when the Chinese character "的" is stored and transmitted in a computer, it is usually expressed in Unicode code U+7684, while the English letter "a" is usually represented by Unicode code U+0061. For generality, character encoding standards cover as many existing characters in the world as possible and reserve space for new characters that may appear in the future, such as Unicode (Universal Coded Character Set, UCS), Universal Character Set (UCS-2 and UCS-4), etc.

[0086] Taking Unicode as an example, it almost covers all the general artificial characters that have ever existed in the world. Its basic working method is to arrange a unique encoding code position for each included character, so that the computer can uniformly present and process text (characters). Currently, the characters in most writing systems in the world have been sorted and encoded, covering computer control characters, various phonetic letters, ideographic characters such as Chinese characters, mathematical symbols, punctuation marks, square elements, geometric shapes, Braille patterns, character structure relationships, special characters, etc., and more than 140,000 characters have been included. Characters in the same language or language family are distributed in one or more code segments of the Unicode code table in sections. For example, English letters are distributed in the U+0041~U+007A segment of Plane 0; the code positions before U+024F are mainly distributed with Latin letters; Greek and Coptic are located in the U+0370~U+03FF code segment. In addition to the visual glyphs, encoding methods, and standard character encoding data, Unicode also includes a database of character characteristics (such as uppercase and lowercase letters), writing directions, splitting standards, etc.

[0087] Computer glyphs (hereinafter referred to as glyphs) are usually included in font files to express the appearance of each character. The most commonly used glyph expression method is outline glyphs, also known as stroked glyphs. Such glyphs use Bezier curves to describe the character outline and can be enlarged or reduced through simple mathematical transformations. The presentation of digital text content on devices such as monitors and printers depends on the font file. Specifically for a certain character, it depends on the glyph file corresponding to the Unicode code position in the font file. For example, when the computer reads Unicode code U+7684 and wants to display it, it will access the corresponding font glyph library, find the specific glyph vector file (early it was a dot matrix file) corresponding to Unicode code U+7684 in the font, and render it according to auxiliary parameters such as font size and then present it.

[0088] Because many writing systems worldwide evolved from the same ancient script, the phenomenon of different languages ​​sharing the same letters is quite common, especially in Latin-based languages ​​such as English, Russian, and Greek. For example, the English capital letter "A" (Unicode 0041) is visually identical to the Greek capital letter "ALPHA" (Unicode 0391) and the Cyrillic capital letter "A" (Unicode 0410), but to a computer, they are three completely different characters, with three different encodings and three different meanings.

[0089] In this embodiment, Unicode is used as an example to generate a mapping table of homographs for characters that are the same or similar in glyphs across multiple languages. Table 1 below shows the homograph mapping table for English characters.

[0090]

[0091]

[0092] Table 1. Homograph Map of English Characters

[0093] Table 1 provides a set of interchangeable or similar glyphs for basic English characters. The original characters are English characters, including the basic alphabet (26 lowercase and 26 uppercase) and other commonly used symbols such as punctuation. The target characters can be from any language other than English, such as Greek, Latin, or Cyrillic. For example, the interchangeable or similar glyphs in other languages ​​corresponding to the lowercase letter 'a' in English include two: the lowercase Latin letter 'ALPHA' and the lowercase Cyrillic letter 'A'. It should be noted that the above character mapping table is not the only option; it is merely an example to more clearly illustrate the underlying principle of this step. Other forms of mapping tables are also acceptable.

[0094] In one example, step S101 may include the following sub-steps:

[0095] Step S1011: Obtain the character set of the target language.

[0096] The character set of a target language may include the basic alphabet and other commonly used characters such as punctuation, referred to as the primitive character set. Taking English as an example, the primitive character set mainly includes: uppercase letters A to Z, lowercase letters a to z, the English comma ",", the English period ".", the English question mark "?", etc.

[0097] Step S1013: For the characters in the character set of the target language, obtain other language characters that are the same or similar in glyph to the characters, and form the character correspondence relationship.

[0098] In practice, for each character in the original character set, the system can iterate through the characters one by one or selectively search for characters that are the same as or similar to them in the Unicode code table. All characters that are similar to or similar to the characters in the original character set are collectively referred to as the target character set. As shown in Table 1, the original character is "lowercase English letter 'a'", and its corresponding target characters include: lowercase Latin letter ALPHA and lowercase Cyrillic letter A.

[0099] In one example, step S101 may further include the following steps: obtaining the character frequency of a character in the character set of the target language; and executing step S1013 for characters whose character frequencies meet certain conditions, such as executing step S1013 for high-frequency characters whose character frequencies are greater than or equal to a character frequency threshold. In specific implementation, the character frequency of the character in the target usage scope or target context can be obtained. If the character frequency cannot be obtained, publicly available general character frequency statistics can be obtained. This processing method allows for the execution of step S1013 only for a subset of high-frequency characters when resources (computing resources, storage resources) are limited, or the process can stop after reaching a certain cumulative character frequency. Therefore, it can effectively reduce the computational load of finding homographs, thereby saving resources.

[0100] In one example, step S101 may further include the following steps: obtaining the priority of a character in the character set of the target language; and executing step S1013 for characters whose priorities meet the conditions, such as executing step S1013 only for characters with first-level priority. This processing method allows step S1013 to be executed only for characters with a subset of priorities when resources are limited, thus effectively reducing the computational load of finding homographs and saving resources.

[0101] The homograph character mapping table includes multiple homograph character mapping relationships. Each mapping relationship includes the original character. If the original character has a corresponding character in the target character set, the mapping relationship also includes the target character. The original character can correspond to characters in multiple target character sets.

[0102] In one example, the method may further include the following steps: determining the glyph matching degree between target language characters and characters in other languages; and converting the target language characters in the content into characters in other languages ​​with the same or similar glyphs based on the glyph matching degree, the character correspondence, and the source information of the content. This processing method effectively converts the target language characters in the content into characters in other languages ​​with the same or similar glyphs that meet the glyph matching requirements. Therefore, it can effectively improve the user-friendliness of the content after the homograph conversion, thereby enhancing the user experience.

[0103] In practice, the matching level of all target characters can be evaluated based on the similarity between the original character and the target character in the homograph mapping table. For example, if the target language is English, the original character "a" (Unicode code 0061) perfectly matches the Cyrillic lowercase letter A: "a" (Unicode code 0430), and they are interchangeable in most scenarios; it is also quite similar to the Latin lowercase letter ALPHA: "ɑ" (Unicode code 0251), and they can be used interchangeably in some situations. The determination of the matching level is optional and not mandatory, nor is it immutable.

[0104] In one example, the method may further include the following steps: determining auxiliary information of the character; and converting the target language character in the content into other language characters with the same or similar glyphs based on the auxiliary information, the character correspondence, and the source information of the content. The auxiliary information may include full-width / half-width characters, character width, etc., depending on the specific circumstances. This processing method allows the target language character in the content to be converted into other language characters with the same or similar glyphs based on the auxiliary information, such as requiring both the converted and unconverted characters to be full-width characters, or requiring the converted and unconverted characters to have similar character widths, etc. Therefore, it can effectively improve the user-friendliness of the content after homograph conversion, thereby enhancing the user experience.

[0105] In one example, step S101 can be implemented as follows: obtaining font information of the target environment; obtaining the languages ​​supported by the font; and constructing the character mapping relationship based on the languages ​​supported by the font. This processing method ensures that the content converted from homographs according to the character mapping relationship can be correctly displayed in the target environment. Therefore, it can effectively improve the quality of the content after homograph conversion, thereby enhancing the user experience.

[0106] In practice, code point compatibility testing should be performed for the target environment. In some cases, some characters in the homograph map may be uncommon, such as the Armenian capital letter OH: "O" (Unicode code 0555), which is not supported by Microsoft YaHei font. Therefore, before constructing the homograph map, font compatibility testing should be performed for the target environment to remove unsupported characters. If the target environment is unclear, a conservative compatibility test can be performed for a general environment to obtain a universal homograph map with compatibility.

[0107] In practice, a homograph library can be obtained based on a homograph mapping table, matching level evaluation results, compatibility test results, and other auxiliary information. Based on the homograph library, some or all of the target language characters in the target language content can be converted into other language characters with the same or similar glyphs.

[0108] Step S103: Generate content in the target language using a content generation algorithm.

[0109] The content generation algorithm can be a content generation model trained using machine learning methods, such as a Generative Pre-trained Transformer (GPT) model trained using a large language model and reinforcement learning, which has powerful text content generation capabilities. In practice, other forms of content generation algorithms can also be used to automatically generate content. Since content generation algorithms are relatively mature existing technologies, they will not be elaborated upon here.

[0110] In practice, step S103 can be implemented as follows: the client sends a content generation request to the server; the server responds to the request and generates content in the target language through the content generation model.

[0111] With the widespread application of content generation algorithms, numerous hidden dangers and security issues have quickly become apparent. Computer-generated content is beginning to appear in many places where it shouldn't be. A typical scenario is its impact on the education system: students can use algorithm-generated content to quickly complete assignments and assessments, or be misled by questionable generated content. More broadly, in an increasing number of scenarios, it is essential to determine whether the content in the target language originates from a content generation algorithm and to take appropriate measures accordingly.

[0112] Furthermore, while deep pre-trained transformation models can quickly and continuously generate large amounts of text content, the current model-generated content is based on learning and reorganizing past knowledge. This results in inconsistent content quality, unreliable accuracy, redundancy, repetition, and rigidity in the output. This model-generated content is continuously being generated and entering human knowledge bases. This phenomenon will lead to problems such as the pollution of human knowledge bases by intelligently generated content and severe dilution of the knowledge base.

[0113] After constructing the above-mentioned homograph correspondence and generating the target language content through the algorithm, we can proceed to the next step of embedding watermarks into the generated content. The embedded watermarks can at least determine whether the content was automatically generated by the algorithm or created by humans.

[0114] Step S105: Based on the character correspondence and the source information of the content, convert some or all of the target language characters in the content into other language characters with the same or similar glyphs.

[0115] After the target language content is generated, content source information (watermark information) is added to it in the form of a watermark. According to the character correspondence obtained in step S101, the characters in the target language content can be divided into three categories: 1) Watermarked characters, that is, homographs that can be substituted for each other in the character correspondence, such as the uppercase letter "A" in Table 1; 2) Non-watermarked characters, that is, characters that are in the character correspondence but do not have corresponding homographs, such as the lowercase letter "t" in Table 1; 3) Other characters, that is, other characters outside the character correspondence.

[0116] In the process of adding watermarks to the target language content, different operations are performed on the above characters: the watermark character has two responsibilities, one is to act as a carrier for hiding the watermark information, that is, to modulate the watermark character according to the watermark information content, and the other is to act as a foreign character (characters from other languages) in the target language, thus destroying the semantics and encoding structure of the original target language content; no operation is performed on the other two types of characters (non-watermark characters and other characters).

[0117] In one example, the source information of the content is that the content is generated by an algorithm; step S105 can be implemented as follows: based on the character correspondence and the source information of the content, convert some or all of the target language characters into any other language character with the same or similar glyphs. For example, randomly select one of the homographs in other languages ​​corresponding to the target language character to replace the target language character. Using this processing method is equivalent to adding a 1-bit identifier to each watermark character that undergoes homograph conversion. That is, although the content of the target language does not change visually to the human eye, or only undergoes very subtle and imperceptible changes, its underlying encoding has been mapped to characters in other languages. The other language characters in the content after homograph conversion become the carrier of the "content generated by the algorithm watermark," and the "content generated by the algorithm watermark" can be extracted based on the other language characters in the content after homograph conversion.

[0118] For example, suppose the original English content is "eat," with a Unicode encoding of 0065-0061-0074. After adding the watermark, it becomes "еat." The original English character "e" is replaced with the Cyrillic lowercase letter "IE," retaining the shape of "e," but its encoding becomes Unicode 0435. Similarly, the original English character "a" is replaced with the Cyrillic lowercase letter "A," retaining the shape of "a," but its encoding becomes Unicode 0430. The original letter "t" is not a watermark character and therefore remains unchanged. Thus, after adding the watermark, the Unicode encoding of the content "еat" becomes 0435-0430-0074, which still appears as the original word to the human eye. Furthermore, for a computer, this word has lost its original semantic meaning; it is unrecognizable when searching the original English content. This approach ensures that automatically generated content cannot be recognized or retrieved, thus isolating it from human-created knowledge bases.

[0119] In another example, the source information of the content is intellectual property information, such as the owner of the model-generated content being "Zhang San"; the method may also include the following steps: constructing a correspondence between homographs and binary numbers corresponding to characters in the target language based on the character correspondence; and converting the intellectual property information into a watermark bitstream. In specific implementation, the process of encoding the intellectual property information into a watermark bitstream can employ any existing encoding method, and error correction codes, check codes, etc., can also be added, which will not be elaborated here. Accordingly, step S105 may include the following sub-steps: 1) If the target language character has homographs, then according to the number of homographs of the target language character, obtain the number of binary bits corresponding to the target language character, also known as the watermark capacity; 2) According to the number of binary bits, read the watermark bit position corresponding to the number of bits from the watermark bit stream; 3) According to the correspondence between the homographs and the binary numbers, obtain the target homograph corresponding to the watermark bit position corresponding to the number of bits; 4) If the target homograph is another language character with the same or similar glyph as the target language character, then replace the target language character with the target homograph. This processing method is equivalent to adding a partial intellectual property watermark to each watermark character that undergoes homograph conversion. That is, although the content of the target language does not change visually to the human eye, or only undergoes very subtle and imperceptible changes, its underlying encoding has been mapped to characters in other languages. The other language characters in the content after homograph conversion become the carrier of "the content is generated by a certain intellectual property party's watermark". Intellectual property information can be extracted from the other language characters in the content after homograph conversion.

[0120] The correspondence between homographs and binary numbers corresponding to characters in the target language can be used to add watermark information, such as intellectual property watermarks, to the content of the target language. These homographs include the target language characters themselves, as well as characters from other languages ​​that are identical or similar in font to the target language characters. Each homograph corresponds to a different binary number, and the number of bits in each binary number is the same. For example, the lowercase letter "o" in Table 1 has the same or similar glyphs as characters in three other languages. Therefore, the homographs corresponding to the lowercase letter "o" include the lowercase letter "o" itself, as well as characters from the other three languages. These four homographs correspond to different binary numbers, each with two bits. Table 2 shows the correspondence between homographs and binary numbers corresponding to the lowercase letter "o".

[0121] 00 O, the lowercase Latin letter O 004F 01 O, the Mystery of the Greek Letter Cloning 039F 10 О, Cyrillic lowercase O 041E 11 O, Armenian lowercase OH 0555

[0122] Table 2. Correspondence between homographs of the lowercase English letter "o" and their binary representations.

[0123] As shown in Table 2, the binary number corresponding to the lowercase English letter "o" is 00, the binary number corresponding to the Greek letter Omkol is 01, the binary number corresponding to the lowercase Cyrillic letter O is 10, and the binary number corresponding to the lowercase Armenian letter OH is 11.

[0124] In specific implementation, based on the character correspondence, a correspondence between homographs and binary numbers corresponding to the target language characters is constructed. This can be achieved as follows: Based on the character correspondence, the number of homographs of the target language characters is obtained; based on the number of homographs, the number of bits in the binary number is determined; based on the number of bits in the binary number, a correspondence between the homographs and the binary number is constructed. Using this method, the number of bits in the binary number corresponding to each target language character can be calculated based on the number of homographs. For example, the number of bits in the binary number can be log₂N, where N represents the number of homographs. For instance, the lowercase English letter "o" in Table 1 has the same or similar glyphs as characters in three other languages, meaning the number of homographs is 4. The number of bits in the binary number corresponding to the lowercase English letter "o" is log₂₄ = 2. Based on the 2 bits in the binary number, the correspondence in Table 2 can be constructed, where each homograph corresponds to a different two-bit binary number.

[0125] In specific implementation, based on the character correspondence, a correspondence between homographs of the target language characters and their binary numbers is constructed. This can be achieved by using encryption to construct the correspondence between homographs of the target language characters and their binary numbers based on the character correspondence. Taking Table 2 above as an example, if encryption is used, a correspondence between homographs and their binary numbers as shown in Table 3 may be constructed.

[0126] 00 O, Armenian lowercase OH 0555 01 О, Cyrillic lowercase O 041E 10 O, the Mystery of the Greek Letter Cloning 039F 11 O, the lowercase Latin letter O 004F

[0127] Table 3. Correspondence between homographs of the encrypted lowercase English letter "o" and their binary representations.

[0128] Comparing the mapping relationships in Tables 2 and 3 reveals that the same binary number before and after encryption corresponds to different homographs. This avoids relying on the correspondence of characters with similar or identical glyphs across multiple languages, allowing direct deduction of the binary number corresponding to each homograph. When using the mapping relationship between encrypted homographs and binary numbers to add watermark information to target language content, the security of the watermark information can be effectively improved. It is important to emphasize that the above tables are merely examples; in practical applications, the mapping relationship between homographs and binary numbers can be designed according to specific circumstances and requirements.

[0129] In practice, the target language content can be processed character by character in sequence as follows: If the character does not belong to the watermark bit segment, skip it; if the character belongs to the watermark bit segment, perform the following operations: First, obtain the correspondence between the homographs and binary numbers of the watermark character segment and the number of binary digits; then, based on the number of binary digits of the current watermark character segment, read the same length of watermark bits to be embedded from the watermark bit stream. Based on the read watermark bits, select different glyphs and code positions from the correspondence between the homographs and binary numbers of the watermark character segment. For example, if the watermark character segment "lowercase letter O" has two watermark bits, code position modulation can be performed according to Table 2 above. After completing the embedding of the required watermark code characters or all watermark code characters, output the watermarked content corresponding to the target language content.

[0130] In one example, the method may further include the following steps: storing the content converted from homographs corresponding to the content in the target language into a knowledge base, which may also include human-created content. The knowledge base performs content retrieval based on the character encoding corresponding to the content retrieval conditions. For example, if the content retrieval condition is that the content field contains "eat," the corresponding character encoding is the encoding of English characters. When retrieving content in the knowledge base based on the English character encoding, it will be impossible to retrieve corresponding content composed of characters from other languages ​​that are identical or similar in glyph. Since the content converted from homographs ensures that the human visual effect remains unchanged, while destroying the inherent encoding and computer semantics of generative content, it can effectively isolate the content converted from homographs generated by the model from human-created content, avoiding contamination of the human knowledge base.

[0131] Figure 2This illustration shows an application scenario of the method provided in this embodiment. In this embodiment, a first client sends a content generation request to the server, such as a dialogue request, a request to obtain an answer to a target question, a request to obtain creative content for a target creative question, a programming request, etc. The server responds to the content generation request and generates content in the target language using a content generation algorithm. Prior to this, a homograph library needs to be constructed, as shown in Table 1. Then, based on the character correspondence and the source information of the content, some or all of the target language characters in the content are converted into characters of other languages ​​with the same or similar glyphs. The server returns the watermarked content to the first client. The user of the first client can provide this content to the user of a second client. The user of the second client, to know the source of the content, can send a content generation method detection request to the server. The server responds to this request, obtains the content generation method information (e.g., algorithm-generated or human-created) using the generative content detection algorithm provided in this embodiment, and sends the content generation method information back to the second client.

[0132] like Figure 2 As shown, assuming the content source information includes the intellectual property information of the content generator, also known as the plaintext information of the watermark, such as if the target language content is generated by an algorithm provided by Zhang San, the character encoding of the intellectual property information "Zhang San" (U+5F20 U+0033) can be obtained. This character encoding is then converted into binary numbers to form the watermark bitstream (010111111…). For the content generated by the model: Oct. 3 rd The generated content (after homograph conversion) is my mother's birthday. While visually identical to the original generated content, the homograph conversion disrupts the inherent encoding and computer semantics of the generated content. Specifically, for the original character "O(U+004F)", the target character 1: O(U+039F) is determined by the watermark bit (the aforementioned binary number) "01" as "same glyph, different character encoding"; for the original character "c(U+0063)", the target character 1: c(U+0441) is determined by the subsequent watermark bit "01" as "same glyph, different character encoding"; for the original characters "t、3、r", no watermark is embedded, and the character encoding remains unchanged; for the original character "d(U+0064)", the target character 1: d(U+0501) is determined by the subsequent watermark bit "1" as "similar glyph, different character encoding".

[0133] like Figure 2 As shown, assume the content source information is that the content is generated by an algorithm, and the watermark is a 1-bit robust identifier. For the content generated by the model: Oct. 3 rdThe generated content (after homograph conversion) is my mother's birthday and has a watermark. Visually, it is almost identical to the original generated content, but the homograph conversion disrupts the inherent encoding and computer semantics of the generated content. Specifically, for the original character "O(U+004F)", any target character with the same glyph but different character encoding is selected based on a 1-bit robust identifier; for the original character "c(U+0063)", any target character with the same glyph but different character encoding is selected based on a 1-bit robust identifier; for the original characters "t, 3, r", no identifier information is embedded, and the character encoding remains unchanged; for the original character "d(U+0064)", any target character with the same glyph but different character encoding is selected based on a 1-bit robust identifier.

[0134] As can be seen from the above embodiments, the content generation method provided in this application constructs character correspondences for multiple languages ​​that are identical or similar in glyphs; generates content in the target language using a content generation algorithm; and converts some or all of the target language characters in the content into characters of other languages ​​with identical or similar glyphs based on the character correspondences and the source information of the content. This method of adding a generative content watermark to the generative content through homograph conversion effectively increases the watermark capacity in the generative content because the watermark is at the character level and has a large coverage for alphabetic characters. Furthermore, the watermarked content, composed of character encodings from multiple languages, is significantly different from human-created content. Under this premise, even if the text volume of the generative content is small, it is possible to accurately determine whether the content has been watermarked to identify the content generation method, i.e., whether it was generated by a content generation algorithm. This allows for effective identification of the content generation method with only a small amount of content, preparing for a significant improvement in the accuracy of content generation method identification when the content volume is small. Furthermore, because the determination of watermark presence can be accurate down to the word and phrase level, it can effectively detect the generation method even after subsequent editing operations such as insertion, deletion, and modification of the generative content. It exhibits high robustness to editing operations not present in the original generative content. Moreover, by using homograph modulation, the inherent encoding and computer semantics of the generative content are disrupted while maintaining the human visual effect. This effectively isolates model-generated content from human-created content, preventing contamination of the human knowledge base.

[0135] Second Embodiment

[0136] In the above embodiments, a content generation method is provided. Correspondingly, this application also provides a content generation apparatus. This apparatus corresponds to the embodiments of the method described above. Since the apparatus embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. The apparatus embodiments described below are merely illustrative.

[0137] This application also provides a content generation apparatus, including: a homograph correspondence construction unit, an algorithm-generated content unit, and a homograph conversion unit.

[0138] The homograph correspondence construction unit is used to construct the correspondence between characters that are the same or similar in glyphs in multiple languages; the content generation algorithm unit is used to generate content in the target language through a content generation algorithm; and the homograph conversion unit is used to convert some or all of the target language characters in the content into characters of other languages ​​that have the same or similar glyphs, based on the character correspondence and the source information of the content.

[0139] In one example, the source information of the content is that the content is generated by an algorithm; the homograph conversion unit is specifically used to convert some or all of the target language characters into any other language character with the same or similar glyphs according to the character correspondence and the source information of the content.

[0140] In one example, the source information of the content is intellectual property information; the device further includes: a homograph-to-binary number correspondence construction unit, used to construct a correspondence between homographs and binary numbers corresponding to target language characters based on the character correspondence; a watermark bitstream acquisition unit, used to convert the intellectual property information into a watermark bitstream; the homograph conversion unit is specifically used to, if the target language character has homographs, obtain the number of binary numbers corresponding to the target language character based on the number of homographs of the target language character; read the corresponding number of watermark bits from the watermark bitstream based on the number of binary numbers; obtain the target homograph corresponding to the corresponding number of watermark bits based on the correspondence between homographs and binary numbers; if the target homograph is a character from another language with the same or similar glyphs as the target language character, replace the target language character with the target homograph.

[0141] In one example, the homograph-to-binary-number correspondence construction unit is specifically used to obtain the number of homographs of the target language character based on the character correspondence; determine the number of bits in the binary number based on the number of homographs; and construct the correspondence between the homographs and the binary number based on the number of bits in the binary number.

[0142] In one example, the homograph-to-binary-number correspondence construction unit is specifically used to construct, through encryption, a correspondence between homographs and binary numbers corresponding to characters in the target language, based on the character correspondence.

[0143] In one example, the device further includes a content storage unit for storing content converted from homographs corresponding to the content of the target language into a knowledge base, the knowledge base also including human-created content, and the knowledge base performing content retrieval based on the character encoding corresponding to the content retrieval conditions.

[0144] In one example, the device further includes: the homograph character correspondence construction unit, specifically used to obtain font information of the target environment; obtain the languages ​​supported by the font; and construct the character correspondence based on the languages ​​supported by the font.

[0145] Third Embodiment

[0146] In the above embodiments, a content generation method is provided. Correspondingly, this application also provides a generative content detection method.

[0147] This application also provides a generative content detection method, including the following steps:

[0148] Step S301: Obtain the content of the target language.

[0149] In specific implementation, the execution subject of the method is the server; accordingly, step S301 can be implemented in the following way: the client sends a content generation method detection request to the server; the server obtains the content of the target language from the request.

[0150] Step S303: Obtain the correspondence between characters that are the same or similar in glyphs in multiple languages.

[0151] Once the target language content of the content to be identified is obtained, to determine whether it was automatically generated by a content generation algorithm, or to extract hidden intellectual property watermark information, text watermark extraction processing (equivalent to decoding) is required, which corresponds to the text watermarking process (equivalent to encoding) in the content generation process. The target language content to be decoded may also contain the three types of characters mentioned above: watermarked characters, non-watermarked characters, and other characters. Non-watermarked characters may be those that appear after manual text editing (such as adding, deleting, modifying, or splicing) of the content converted from homographs obtained in Example 1. The decoding and watermark extraction operations for the target language content are the reverse of the steps in Example 1. Therefore, it is also necessary to obtain the correspondence between characters that are identical or similar in glyphs across multiple languages.

[0152] Step S305: Based on the character correspondence, convert other language characters in the content of the target language into target language characters with the same or similar glyphs.

[0153] For example, if the target language is English and the content is "еat", the character "е" that looks like the English letter "e" is actually the Cyrillic lowercase letter IE. After homographing the content, the result is "eat".

[0154] Step S307: If the content fragment composed of multiple target language characters after conversion has semantic information, then the content of the target language is determined to be generative content.

[0155] For example, the English word "eat" has semantic meaning, therefore it is generative content.

[0156] In one example, step S307 can be implemented as follows: obtain the number of characters converted from homographs; if the content fragment composed of multiple target language characters after conversion has semantic information and the number of characters is greater than a first threshold, then the content of the target language is determined to be generative content. Using this processing method, the generation method of the target language content can be quickly determined.

[0157] In practice, the watermark counter can be set to 0. For each watermark character in the target language content, it is sequentially observed whether it belongs to the original character (target language character) or the target character (other language character) in the same character library. For each watermark character observed from the target character region, the watermark counter is incremented by 1. Simultaneously, the current watermark character is converted back to its corresponding original character, and it is observed whether contextual semantics are formed. If the following two conditions are met: 1. When more than N (first threshold) watermark characters from the target character region are observed; 2. The corresponding watermark character, after being converted back to its original character, forms contextual semantics, then it can be determined that the content contains an identifier watermark, meaning the content was automatically generated by the content generation algorithm. If the target language content is mostly composed of original characters, and the watermark count does not reach N, then it can be determined that the content does not contain an identifier watermark, meaning the content was created manually. If the watermark count reaches N, but the corresponding watermark character, after being converted back to its original character, does not form contextual semantics, then it can be determined that the content does not contain an identifier watermark, meaning the content was created manually. For example, "abd" and "ickd" are target characters, but they have no semantic meaning after being converted to their original characters, indicating that this content is manually created. In practice, the first threshold N can be set according to the usage environment. Determining the contextual semantics can be implemented through algorithms or manual methods.

[0158] In one example, the method may further include the following steps:

[0159] Step S401: Obtain the correspondence between homographs of the target language characters and their binary numbers.

[0160] This step corresponds to the steps in Embodiment 1, and will not be repeated here.

[0161] Step S403: Based on the correspondence between the homographs and binary numbers, obtain the binary number corresponding to the homograph in the content of the target language.

[0162] Homographs in the target language content are also called watermark characters. These characters have the same or similar character correspondences in the glyphs of the various languages. Based on the correspondence between the homograph and its corresponding binary number, the binary number corresponding to the homograph in the target language content can be obtained. For example, if the target language content includes the Greek character 'o', which is a watermark character, and there are four homographs, then according to Table 2, the binary number corresponding to the Greek character 'o' is "01".

[0163] Step S405: Add the binary number corresponding to the homograph to the watermark bitstream.

[0164] Step S407: After extracting the watermark information, the watermark bitstream is converted into intellectual property information.

[0165] In practice, the watermark bitstream can be initialized first. Characters from the target language content are read sequentially. If a character is not a watermark character, it is skipped. If a character is a watermark character, the following operations are performed: 1) The number of binary digits for the watermark character can be determined from auxiliary information, i.e., the watermark information embedding capacity, or simply watermark capacity; 2) Based on the correspondence between homographs and binary numbers (as shown in Table 2), the binary number corresponding to the current homograph is obtained; 3) The extracted binary numbers corresponding to the homographs are sequentially added to the watermark information stream; 4) After all watermark information is extracted, the watermark information stream is converted into character encoding, which displays the intellectual property information in text form.

[0166] As can be seen from the above embodiments, the content generation method provided in this application obtains the correspondence between characters that are identical or similar in glyphs in multiple languages; obtains the content of the target language; converts other language characters in the content of the target language into target language characters with identical or similar glyphs based on the character correspondence; if the content fragment composed of multiple converted target language characters has semantic information, then the content of the target language is determined to be generative content. This processing method allows for the rapid determination of the content generation method by converting homographs into target language characters and identifying the generation method of the target language content based on whether multiple target language characters have contextual semantics.

[0167] Fourth embodiment

[0168] In the above embodiments, a generative content detection method is provided. Correspondingly, this application also provides a generative content detection apparatus. This apparatus corresponds to the embodiments of the above method. Since the apparatus embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. The apparatus embodiments described below are merely illustrative.

[0169] This application also provides a generative content detection device, including: a homograph correspondence acquisition unit, a content acquisition unit, a homograph conversion unit, and a determination unit.

[0170] The unit for obtaining homograph correspondence is used to obtain the correspondence between characters that are the same or similar in shape in multiple languages; the content acquisition unit is used to obtain the content of the target language; the homograph conversion unit is used to convert characters of other languages ​​in the content of the target language into characters of the target language with the same or similar shape according to the character correspondence; and the determination unit is used to determine that the content of the target language is generative content if the content fragment composed of multiple converted target language characters has semantic information.

[0171] In one example, the determination unit is specifically used to obtain the number of characters converted from homographs; if the content fragment composed of multiple target language characters after conversion has semantic information and the number of characters is greater than a first threshold, then the content of the target language is determined to be generative content.

[0172] In one example, the device further includes: a unit for obtaining the correspondence between homographs and binary numbers, a unit for obtaining the binary numbers of homographs, a unit for adding binary numbers, and a unit for obtaining intellectual property information.

[0173] The unit for obtaining the correspondence between homographs and binary numbers is used to obtain the correspondence between homographs and binary numbers of characters in the target language; the unit for obtaining the binary number of homographs is used to obtain the binary number corresponding to the homograph in the content of the target language according to the correspondence between homographs and binary numbers; the unit for adding binary numbers is used to add the binary number corresponding to the homograph to the watermark bitstream; and the unit for obtaining intellectual property information is used to convert the watermark bitstream into intellectual property information after the watermark information extraction is completed.

[0174] Fifth embodiment

[0175] In the above embodiments, a generative content detection method is provided. Correspondingly, this application also provides a generative content detection method. The method in this embodiment corresponds to Embodiment 3 of the above method. Since this embodiment is basically similar to Embodiment 3, it is described simply, and relevant parts can be referred to in the description of Embodiment 3. The method embodiments described below are merely illustrative.

[0176] This application also provides a generative content detection method, including the following steps:

[0177] Step S501: Obtain the content of the target language.

[0178] Step S503: Obtain the correspondence between characters that are the same or similar in glyphs in multiple languages.

[0179] Step S505: Based on the character correspondence, convert other language characters in the content of the target language into target language characters with the same or similar glyphs.

[0180] Step S507: Set the characters in the content that have undergone homograph conversion to the corresponding first flag, set the characters in the content that have not undergone homograph conversion to the corresponding second flag, set the non-homograph characters in the content to the corresponding third flag, and obtain the character flag string corresponding to the content based on the position of the character in the content and the flag corresponding to the character.

[0181] The method provided in this application can locate which segment of information in the target language content is automatically generated by an algorithm and which segment is created by humans. Homographs, also known as watermark characters, are identified by setting the watermark characters that undergo homograph conversion in the content as the corresponding first flag, the watermark characters that do not undergo homograph conversion as the corresponding second flag, and the non-watermark characters as the corresponding third flag. Based on the position of the character in the content and the corresponding flag, a character flag string corresponding to the target language content is obtained.

[0182] Step S509: For the second flag segment in the character flag string, if the length of the second flag segment is less than the second threshold and the character segment corresponding to the flag segment adjacent to the second flag segment has semantic information, then the second flag segment and the character segment corresponding to the adjacent flag segment are used as generative content.

[0183] In one example, the device may further include the following step: if the length of the second flag bit segment is greater than or equal to the second threshold and the length is less than the third threshold, then the character segment corresponding to the second flag bit segment is regarded as content of unknown origin.

[0184] In one example, the device may further include the following step: if the length of the second flag segment is greater than or equal to a third threshold, then the character segment corresponding to the second flag segment is regarded as manually created content.

[0185] The character flag string may include one or more second flag segments. In specific implementation, a location information table can be initialized. Characters from the target language content are read sequentially, and the following operations are performed: 1) If the character is a watermark character and comes from the target character region in the homograph character library, a location identifier "1" is added to the location information table to represent the first flag; 2) If the character is a watermark character but comes from the original character region in the homograph character library, a location identifier "0" is added to the location information table to represent the second flag; 3) If the character is not a watermark character, a location identifier "-1" is added to the location information table to represent the third flag. After the location identifiers for all characters in the target language are extracted, the entire location information table content is used to divide the area into watermarked regions (first flag segments) and non-watermarked regions (second flag segments). There are many ways to process the entire location information table. Here is one feasible method as an example: When the distance between two location identifiers (first flag bits) "1" (the length of the flag bits between them, i.e., the length of the second flag bit segment) is less than the second threshold T1, and the character segments corresponding to the surrounding flag bit segments (such as the first flag bit segment) have semantic information, such as 11100001001, if T1 is 5, then the condition is met. The first flag bit segments between these two location identifiers are both marked as watermarked content areas (i.e., generated content). Otherwise, only the content where the identifier "1" is located (such as words or...) is marked. The phrase is marked as a watermarked content area (i.e., generated content). When there is no identifier "1" between three consecutive positioning identifiers "0" (e.g., 11100000000001001), and if T2 is 10, the content corresponding to this consecutive second flag segment is marked as a non-watermarked content area (i.e., manually created content). Second flag segments other than those mentioned above are marked as questionable areas. For example, 11100000001001, with seven consecutive zeros, greater than T1 and less than T2, these seven zeros are questionable areas and can be represented for manual identification of the generation method of the content. Among them, the content corresponding to the watermarked area can be determined to be automatically generated by the content generation algorithm, the content corresponding to the non-watermarked area can be determined to be created by human participation, and the creation source of the questionable area is unknown.

[0186] As can be seen from the above embodiments, the content generation method provided in this application obtains the character correspondences of multiple languages ​​that are identical or similar in glyphs; obtains the content of the target language; converts other language characters in the content of the target language into target language characters with identical or similar glyphs according to the character correspondences; sets the characters in the content that have undergone homograph conversion as corresponding first flags, sets the characters in the content that have not undergone homograph conversion as corresponding second flags, sets the non-homograph characters in the content as third flags; obtains a character flag string corresponding to the content according to the position of the character in the content and the flags corresponding to the character; for the second flag segment in the character flag string, if the length of the second flag segment is less than a second threshold and the character segment corresponding to the flag segment adjacent to the second flag segment has semantic information, then the second flag segment and the character segment corresponding to the adjacent flag segment are used as generative content. This processing method allows for the identification of which part of the target language content is automatically generated by the algorithm and which part is created by humans, thus effectively improving the accuracy of content generation method recognition.

[0187] Sixth Embodiment

[0188] In the above embodiments, a generative content detection method is provided. Correspondingly, this application also provides a generative content detection apparatus. This apparatus corresponds to the embodiments of the above method. Since the apparatus embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. The apparatus embodiments described below are merely illustrative.

[0189] This application also provides a generative content detection device, including: a homograph correspondence acquisition unit, a content acquisition unit, a homograph conversion unit, a flag determination unit, and a first positioning unit.

[0190] The system includes: a homograph correspondence acquisition unit for acquiring the correspondence of characters that are identical or similar in glyphs across multiple languages; a content acquisition unit for acquiring the content of a target language; a homograph conversion unit for converting characters from other languages ​​in the content of the target language into characters of the target language that have the same or similar glyphs, based on the character correspondence; a flag determination unit for setting characters in the content that have undergone homograph conversion as corresponding first flags, characters in the content that have not undergone homograph conversion as corresponding second flags, and non-homograph characters in the content as third flags, and acquiring a character flag string corresponding to the content based on the position of the character in the content and the flags corresponding to the character; and a first positioning unit for determining, for a second flag segment in the character flag string, if the length of the second flag segment is less than a second threshold and the character segment corresponding to the flag segment adjacent to the second flag segment has semantic information, then taking the second flag segment and the character segment corresponding to the adjacent flag segment as generative content.

[0191] In one example, the device further includes a second positioning unit, configured to, if the length of the second flag bit segment is greater than or equal to the second threshold and the length is less than a third threshold, then classify the character segment corresponding to the second flag bit segment as content of unknown origin.

[0192] In one example, the device further includes a third positioning unit, configured to, if the length of the second flag segment is greater than or equal to a third threshold, treat the character segment corresponding to the second flag segment as manually created content.

[0193] Seventh Embodiment

[0194] In the above embodiments, a content generation method is provided. Correspondingly, this application also provides a question-and-answer method. The method in this embodiment corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0195] This application also provides a question-and-answer method, including the following steps:

[0196] Step S701: Construct a correspondence between characters that are the same or similar in glyphs in multiple languages.

[0197] Step S703: Receive a dialogue request sent by the client.

[0198] Step S705: Generate dialogue content using a content generation algorithm based on the dialogue information carried in the request.

[0199] The execution entity of the method provided in this application embodiment can be the server of a question-answering robot. In this case, the content generation algorithm can be a dialogue content generation algorithm based on a question-answering knowledge base.

[0200] Step S707: Based on the character correspondence, convert some or all of the target language characters in the dialogue content into other language characters with the same or similar glyphs.

[0201] Step S709: Send the dialogue content after homograph conversion back to the client.

[0202] Users can view the dialogue content after homograph conversion through the client. If the user later provides the dialogue content to other users, the other users can use the generative content detection method described above to detect that the content comes from the algorithm's automatic generation.

[0203] Eighth embodiment

[0204] In the above embodiments, a content generation method is provided. Correspondingly, this application also provides a question-answering method. The method in this embodiment corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0205] This application also provides a method for answering questions, including the following steps:

[0206] Step S801: Construct a correspondence between characters that are the same or similar in glyphs in multiple languages.

[0207] Step S803: Receive a request from the client to obtain the answer to the target question.

[0208] Step S805: Generate the answer content corresponding to the target question using a content generation algorithm.

[0209] The execution entity of the method provided in this application embodiment can be the server of an automatic question-answering system. In this case, the content generation algorithm can be an automatic question-answering algorithm based on a question knowledge base.

[0210] Step S807: Based on the character correspondence, convert some or all of the target language characters in the answer content into other language characters with the same or similar glyphs.

[0211] Step S809: Send the answer content after homograph conversion back to the client.

[0212] Users can view the answer content after homograph conversion through the client. If the user later provides the answer content to other users, the other users can use the generative content detection method described above to detect that the answer content comes from the algorithm's automatic generation.

[0213] Ninth Embodiment

[0214] In the above embodiments, a content generation method is provided. Correspondingly, this application also provides a creative generation method. The method in this embodiment corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0215] This application also provides a method for generating creative ideas, including the following steps:

[0216] Step S901: Construct a correspondence between characters that are the same or similar in glyphs in multiple languages.

[0217] Step S903: Receive a request from the client to obtain creative content for the target creative problem.

[0218] Step S905: Generate creative content corresponding to the target creative problem using a content generation algorithm.

[0219] The execution entity of the method provided in this application embodiment can be the server of an automatic creative system. In this case, the content generation algorithm can be an automatic creative algorithm based on a creative knowledge base.

[0220] Step S907: Based on the character correspondence, convert some or all of the target language characters in the creative content into other language characters with the same or similar glyphs.

[0221] Step S909: Send the creative content after homograph conversion back to the client.

[0222] Users can view the creative content after homograph conversion through the client. If the user provides the creative content to other users later, the other users can use the generative content detection method described above to detect that the creative content comes from the algorithm's automatic generation.

[0223] Tenth Embodiment

[0224] In the above embodiments, a content generation method is provided. Correspondingly, this application also provides a program creation method. The method in this embodiment corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0225] This application also provides a method for creating a program, including the following steps:

[0226] Step S901: Construct a correspondence between characters that are the same or similar in glyphs in multiple languages.

[0227] Step S903: Receive the programming request sent by the client.

[0228] Step S905: Generate program code using a content generation algorithm based on the programming requirement information carried in the request.

[0229] The execution entity of the method provided in this application embodiment can be the server of an automatic programming system. In this case, the content generation algorithm can be an automatic programming algorithm based on a program knowledge base.

[0230] Step S907: Based on the character correspondence, convert some or all of the target language characters in the program code into other language characters with the same or similar glyphs.

[0231] Step S909: Send the program code after homograph conversion back to the client.

[0232] Users can view the program code after homograph conversion through the client. If the user provides the program code to other users later, the other users can use the generative content detection method described above to detect that the program code comes from an algorithm that automatically generates it.

[0233] Eleventh Embodiment

[0234] In the above embodiments, a content generation method and a generative content detection method are provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above methods. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. The device embodiments described below are merely illustrative.

[0235] The electronic device in this embodiment includes:

[0236] The electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing any of the above methods, and the device is powered on and runs the program of the method through the processor.

[0237] Memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0238] In specific implementations, the electronic device may also include one or more of the following components: a power supply component, an input / output (I / O) interface, and a communication component. The power supply component provides power to various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device. The I / O interface provides an interface between the processor 503 and peripheral interface modules, which may be a keyboard, click wheel, buttons, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and other devices (such as smartphones, tablets, etc.).

[0239] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0240] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

[0241] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0242] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0243] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0244] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A content generation method, characterized in that, include: Construct correspondences between characters that are identical or similar in glyphs across multiple languages; Generate content in the target language using content generation algorithms; Based on the character correspondence and the source information of the content, some or all of the target language characters in the content are converted into other language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the content; the source information of the content is that the content is generated by an algorithm; The content converted from homographs corresponding to the content of the target language is stored in a knowledge base, and the knowledge base performs content retrieval based on the character encoding corresponding to the content retrieval conditions; The character correspondence is also used to convert other language characters in the content of the target language into target language characters with the same or similar glyphs. If the content fragment composed of multiple converted target language characters has semantic information, the content of the target language is determined to be generative content.

2. The method according to claim 1, characterized in that, The step of converting some or all of the target language characters in the content into other language characters with the same or similar glyphs based on the character correspondence and the source information of the content includes: Based on the character correspondence and the source information of the content, some or all of the target language characters are converted into any other language character with the same or similar glyphs.

3. The method according to claim 1, characterized in that, The source information for the content is intellectual property information; The method further includes: Based on the character correspondence, construct the correspondence between homographs and binary numbers corresponding to characters in the target language; The intellectual property information is converted into a watermarked bitstream; The step of converting some or all of the target language characters in the content into other language characters with the same or similar glyphs based on the character correspondence and the source information of the content includes: If the target language character has homographs, then the number of binary digits corresponding to the target language character is obtained based on the number of homographs of the target language character; Based on the number of bits in the binary number, read the watermark bit corresponding to the number of bits from the watermark bit stream; Based on the correspondence between the homograph and the binary number, obtain the target homograph corresponding to the watermark bit of the corresponding number of bits; If the target homograph is a character from another language that has the same or similar glyph as the target language character, then the target language character is replaced with the target homograph.

4. The method according to claim 3, characterized in that, The step of constructing a correspondence between homographs and binary numbers corresponding to characters in the target language based on the character correspondence includes: Based on the character correspondence, obtain the number of homographs of the target language characters; The number of bits in the binary number is determined based on the number of homographs. Based on the number of bits in the binary number, construct the correspondence between the homograph and the binary number.

5. A generative content detection method, characterized in that, include: Obtain content in the target language; To obtain the correspondence between characters that are the same or similar in glyphs in multiple languages; Based on the character correspondence, other language characters in the content of the target language are converted into target language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the content; If the content fragment composed of multiple target language characters after conversion has semantic information, then the content of the target language is determined to be generative content; The target language content is stored in a knowledge base, which performs content retrieval based on the character encoding corresponding to the content retrieval conditions.

6. The method according to claim 5, characterized in that, If the content fragment composed of multiple target language characters after conversion has semantic information, then the content of the target language is determined to be generative content, including: Get the number of characters converted from the same character; If the content fragment composed of multiple target language characters after conversion has semantic information and the number of characters is greater than the first threshold, then the content of the target language is determined to be generative content.

7. The method according to claim 5, characterized in that, Also includes: Obtain the correspondence between homographs and binary numbers of characters in the target language; Based on the correspondence between the homographs and binary numbers, obtain the binary numbers corresponding to the homographs in the content of the target language; Add the binary number corresponding to the homograph to the watermark bitstream; After the watermark information is extracted, the watermark bitstream is converted into intellectual property information.

8. A generative content detection method, characterized in that, include: To obtain the correspondence between characters that are the same or similar in glyphs in multiple languages; Obtain content in the target language; Based on the character correspondence, other language characters in the content of the target language are converted into target language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the content; The characters in the content that undergo homograph conversion are set as the corresponding first flag, the characters in the content that do not undergo homograph conversion are set as the corresponding second flag, and the non-homograph characters in the content are set as the third flag. Based on the position of the character in the content and the flag corresponding to the character, the character flag string corresponding to the content is obtained. For the second flag segment in the character flag string, if the length of the second flag segment is less than the second threshold and the character segment corresponding to the flag segment adjacent to the second flag segment has semantic information, then the second flag segment and the character segment corresponding to the adjacent flag segment are used as generative content. The target language content is stored in a knowledge base, which performs content retrieval based on the character encoding corresponding to the content retrieval conditions.

9. A question-and-answer method, characterized in that, include: Construct correspondences between characters that are identical or similar in glyphs across multiple languages; Receive dialogue requests sent by the client; Based on the dialogue information carried in the request, dialogue content is generated using a content generation algorithm; Based on the character correspondence, some or all of the target language characters in the dialogue content are converted into other language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the dialogue content. The dialogue content after homograph conversion is sent back to the client; The dialogue content after homograph conversion is stored in a knowledge base, which performs content retrieval based on the character encoding corresponding to the content retrieval conditions. The character correspondence is also used to convert other language characters in the dialogue content into target language characters with the same or similar glyphs. If the content fragment composed of multiple target language characters after conversion has semantic information, the dialogue content is determined to be generative content.

10. A method for answering questions, characterized in that, include: Construct correspondences between characters that are identical or similar in glyphs across multiple languages; Receive requests from clients to retrieve answers to specific questions; The answer content corresponding to the target question is generated using a content generation algorithm; Based on the character correspondence, some or all of the target language characters in the answer content are converted into other language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the answer content; The answer content after homograph conversion is sent back to the client; The answer content after homograph conversion is stored in a knowledge base, which performs content retrieval based on the character encoding corresponding to the content retrieval conditions; The character correspondence is also used to convert other language characters in the answer content into target language characters with the same or similar glyphs. If the content fragment composed of multiple target language characters after conversion has semantic information, the answer content is determined to be generative content.

11. A method for generating creative ideas, characterized in that, include: Construct correspondences between characters that are identical or similar in glyphs across multiple languages; Receive requests from clients to retrieve creative content related to a specific creative question; Creative content corresponding to the target creative problem is generated using a content generation algorithm; Based on the character correspondence, some or all of the target language characters in the creative content are converted into other language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the creative content; The creative content, after being converted from homographs, is sent back to the client. Creative content converted from homographs is stored in a knowledge base, which then performs content retrieval based on the character encoding corresponding to the content retrieval criteria. The character correspondence is also used to convert other language characters in the creative content into target language characters with the same or similar glyphs. If the content fragment composed of multiple target language characters after conversion has semantic information, the creative content is determined to be generative content.

12. A method for creating a program, characterized in that, include: Construct correspondences between characters that are identical or similar in glyphs across multiple languages; Receive programming requests sent by the client; Based on the programming requirements information carried in the request, program code is generated using a content generation algorithm; Based on the character correspondence, some or all of the target language characters in the program code are converted into other language characters with the same or similar glyphs; the other language characters are used to disrupt the semantics and encoding structure of the program code; The program code after homograph conversion is sent back to the client; The program code after homograph conversion is stored in a knowledge base, which performs content retrieval based on the character encoding corresponding to the content retrieval conditions. The character correspondence is also used to convert other language characters in the program code into target language characters with the same or similar glyphs. If the content fragment composed of multiple target language characters after conversion has semantic information, the program code is determined to be generative content.

13. An electronic device, characterized in that, include: processor; as well as A memory for storing a program for implementing the method according to any one of claims 1-12, wherein the device is powered on and the program for running the method is executed by the processor.