Semantic content fingerprinting method and apparatus
By analyzing text content using a generative pre-trained transformer (GPT), generating semantic content features, and creating a transform elastic signature (TRS), the problem of low efficiency in existing text similarity detection is solved, thus achieving efficient intellectual property protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
- Filing Date
- 2023-10-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing text similarity detection technologies are inefficient and struggle to effectively address modern plagiarism threats, especially in the text domain. Traditional methods are time-consuming, labor-intensive, and prone to errors, and current technologies are unable to identify the core semantic features of text content.
Generative pre-trained transformers (GPT) are used for document analysis to generate semantic content features (SCF) of text content and create transform elastic signatures (TRS). The core semantic information of the text content is represented by a tree structure to enhance IP protection capabilities.
It improves the readability and anti-plagiarism capabilities of text content, maintains consistency of thought even after text transformation, simplifies text similarity detection, and enhances intellectual property protection.
Smart Images

Figure CN122122580A_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of intellectual property protection, and more specifically, to a semantic content fingerprinting method and apparatus. Background Technology
[0002] In the digital age, copyright protection and content ownership management have evolved alongside the proliferation of online content. Content identification (ID) systems are a prominent example, effectively used to identify and manage copyrighted content on social media platforms like YouTube. Content ID systems utilize digital fingerprinting technology to create unique identifiers (ID files) for copyrighted audio and video materials, storing these identifiers in comprehensive databases such as the Content ID Database. When a video is uploaded to YouTube, it is systematically compared to audio and video files registered in the Content ID Database. If a match is found, it indicates potential copyright infringement. In such cases, content owners have the flexibility to choose to block videos to prevent viewing, track video viewing statistics, or add advertisements to the video and automatically transfer revenue to the content owner. However, the use of Content IDs is subject to specific standards, effectively limiting access to Content IDs for major corporate entities due to operational requirements. While contemporary content recognition and similarity detection technologies have made significant progress in audio and video content, the text domain still faces technical challenges. Plagiarism detection and text similarity detection in written works have not seen the same progress.
[0003] Currently, several text similarity detection methods have been tried, such as manual detection, text-matching software (TMS), and text watermarking. Manual detection is a traditional method for identifying plagiarism in written works. While effective, this method is labor-intensive, error-prone, and time-consuming, often leading to inconsistencies in plagiarism detection methods within organizations. TMS, also known as plagiarism detection software, has become increasingly prevalent. Existing TMS solutions primarily rely on lexical analysis, comparing specific text segments between documents. Existing text recognition technologies are mainly optimized for web searches, supporting sequences of up to 6 to 10 words. Furthermore, text watermarking, as another method of protecting text content, is inherently fragile and easily removed even with subtle text manipulation. Additionally, targeted text transformations involve blurring selected text features, effectively circumventing existing detection technologies. Given the above, existing intellectual property protection methods, especially in the context of modern plagiarism threats in the text domain, clearly require immediate attention. Therefore, in addressing modern plagiarism threats, there is a technical problem of inefficiency in text similarity detection technologies.
[0004] Therefore, in light of the above discussion, it is necessary to address the issues related to traditional text similarity detection techniques. Summary of the Invention
[0005] This invention provides a semantic content fingerprinting method and apparatus. It offers a technical solution to address the inefficiency of existing text similarity detection techniques, thus countering modern plagiarism threats. The objective of this invention is to provide a technical solution that at least partially resolves the problems encountered in the prior art and to provide an improved semantic content fingerprinting method and apparatus.
[0006] The object of this invention is achieved by the technical solutions provided in the appended independent claims. Advantageous embodiments of this invention are further defined in the dependent claims.
[0007] In one aspect, the present invention provides a semantic content fingerprinting method. The method includes: performing literature analysis on text content using a generative pre-trained transformer (GPT) to obtain one or more semantic content features (SCFs) of the text content, wherein each SCF includes one or more semantic facts about the text content, and each semantic fact includes one or more words describing the text content. The method further includes: creating a transformation resilient signature (TRS) of the text content, wherein the TRS includes the one or more SCFs of the text content.
[0008] The disclosed method described above, by supporting advanced IP protection technologies such as creating a TRS (text fingerprint or semantic structure) for text content in the digital age, can prevent text content from being copied and used without authorization. This method eliminates noise in the text content, improving readability. Furthermore, by using the disclosed method, the underlying ideas of the text content remain unchanged even after various text transformations.
[0009] In one implementation, the method further includes: obtaining fingerprint recognition instructions that define a fingerprint structure for the text content according to a literature analysis guide; loading the text content into the generative pre-trained transformer (GPT) for performing literature analysis on the text content; and obtaining semantic facts about the text content by querying the GPT to describe the text content using one or more words according to the fingerprint recognition instructions.
[0010] Using fingerprint recognition instructions to define a fingerprint structure for text content allows you to create a valid and unique TRS for the text content.
[0011] In another implementation, the semantic facts include one or more of the key information, plot points, unique facts, and themes of the text content.
[0012] In another implementation, the TRS has a tree structure, wherein the internal nodes of the tree structure include the SCF, and the leaf nodes of the tree structure include the semantic facts corresponding to the SCF.
[0013] Using a tree structure to represent TRS can more simply describe the structural relationships between one or more SCFs and semantic facts (corresponding to internal nodes and leaf nodes, respectively).
[0014] In another implementation, the method further includes: obtaining a nested SCF that includes one or more other SCFs; the TRS tree structure includes internal nodes, and the internal nodes include the nested SCFs.
[0015] Using nested SCFs in the TRS tree structure simplifies text similarity detection between two text contents.
[0016] In another implementation, the root node of the TRS tree structure includes one or more of the following: the title of the text content, the identifier (ID) of the owner of the text content, and information. Each of the other nodes in the TRS tree structure includes the type, ID, and description of the corresponding SCF or semantic fact of the text content.
[0017] This helps to enhance the uniqueness of the TRS for the text content.
[0018] In another aspect, the present invention provides a semantic content fingerprinting device. The device includes: a fingerprinting engine for performing literal analysis on text content using a generative pre-trained transformer (GPT) to obtain one or more semantic content features (SCFs) of the text content, wherein each SCF includes one or more semantic facts about the text content, and each semantic fact includes one or more words describing the text content. The device also includes a fingerprint combiner for creating a transformation resilient signature (TRS) of the text content, wherein the TRS includes the one or more SCFs of the text content.
[0019] After executing the above method, the above-mentioned device achieves all the advantages and technical effects of the above method.
[0020] In another aspect, the present invention provides a computer program product including instructions. When the computer program product is executed by a computer, the instructions cause the computer to perform the semantic content fingerprinting method described above.
[0021] After executing the above method, the computer (e.g., the processor in the above-described device) achieves all the advantages and technical effects of the above method.
[0022] In another aspect, the present invention provides a computer-readable storage medium including instructions. When executed by a computer, the instructions cause the computer to perform the semantic content fingerprinting method described above.
[0023] After executing the above method, the computer (e.g., the processor in the above-described device) achieves all the advantages and technical effects of the above method.
[0024] It should be understood that all of the above implementation methods can be combined together.
[0025] It should be noted that all devices, elements, circuits, units, and modules described in this application can be implemented in software or hardware elements or any combination thereof. All steps performed by the various entities described in this application and the functions described for performance by the various entities are intended to indicate that each entity is suitable for or used to perform the corresponding steps and functions. Although specific functions or steps performed by external entities are not reflected in the detailed description of the specific elements of the entities performing those steps or functions in the following detailed description of specific embodiments, it will be apparent to those skilled in the art that these methods and functions can be implemented by corresponding software or hardware elements or any combination thereof. It is understood that the features of the invention are readily combined in various ways without departing from the scope of the invention as defined by the appended claims.
[0026] Other aspects, advantages, features and objects of the invention will become apparent from the accompanying drawings and the detailed description of illustrative implementations as explained in conjunction with the following appended claims. Attached Figure Description
[0027] A better understanding of the above-described invention and the following detailed description of illustrative embodiments can be obtained by reading the accompanying drawings. Exemplary structures of the invention are shown in the drawings to illustrate the invention. However, the invention is not limited to the specific methods and tools disclosed herein. Furthermore, those skilled in the art will understand that the drawings are not drawn to scale. Where possible, the same elements are represented by the same numbers.
[0028] The embodiments of the present invention are described below by way of example only with reference to the following accompanying drawings, in which: Figure 1 This is a flowchart of a semantic content fingerprinting method provided in one embodiment of the present invention; Figure 2 This is a block diagram of a semantic content fingerprint recognition device provided in one embodiment of the present invention; Figure 3 This invention illustrates a semantic content fingerprinting system provided by an embodiment of the present invention; Figure 4 A tree structure of a Transformation Resilient Signature (TRS) for text content provided in one embodiment of the present invention is shown. Figure 5 This is a tree structure of text content TRS provided in another embodiment of the present invention; Figure 6 This is an exemplary scenario of a tree structure for the text content TRS provided in another embodiment of the present invention; Figure 7 This is a tree structure of text content TRS provided in another embodiment of the present invention; Figure 8 This is a flowchart of semantic content fingerprint recognition provided in one embodiment of the present invention.
[0029] In the accompanying diagram, underlined numbers indicate the item containing the underlined number or the item adjacent to the underlined number, while ununderlined numbers refer to items identified by lines linking the ununderlined number to the item. When a number is ununderlined and has an associated arrow, the ununderlined number identifies the general item that the arrow points to. Detailed Implementation
[0030] The following detailed description illustrates embodiments of the invention and ways in which these embodiments can be implemented. While some modes of implementing the invention have been disclosed, those skilled in the art will recognize that other embodiments for implementing or practicing the invention may also exist.
[0031] Figure 1 This is a flowchart illustrating a semantic content fingerprinting method according to an embodiment of the present invention. (Reference) Figure 1 The diagram illustrates a semantic content fingerprinting method 100. Method 100 includes steps 102 and 104.
[0032] This invention provides a semantic content fingerprinting method 100. Method 100 is capable of text fingerprinting using semantic information (i.e., information conveyed by a few words, phrases, or short sentences). Text fingerprinting, also known as a digital signature, is a unique identifier belonging to the content owner or content provider. Method 100 performs document analysis using a generative pre-trained transformer to mine one or more key features of the text content. Subsequently, Method 100 generates a semantic structure using one or more key features of the text content. This semantic structure, by providing content owners or content providers with advanced tools and techniques to protect their creative works in the digital age, can prevent text content from being copied and used without authorization. The steps of Method 100 are described below.
[0033] At step 102, method 100 includes: performing a literature analysis on the text content using a generative pre-trained transformer (GPT) to obtain one or more semantic content features (SCFs) of the text content, wherein each SCF includes one or more semantic facts about the text content, and each semantic fact includes one or more words describing the text content. Examples of text content may include, but are not limited to, white papers, case studies, newsletters, social media posts, scriptwriting, etc. GPT is used to generate one or more SCFs of the text content by performing a literature analysis on the text content. One or more SCFs may also be referred to as meaningful text features, which remain unchanged during various text transformations, and the modern concepts (i.e., ideas) behind the text content are not affected. For example, one or more SCFs may include plot, genre, rhythm (i.e., poetry), one or more main ideas, one or more key information, interesting points, unusual points, important points, context (e.g., background, time, geographical setting), author's point of view, height (tone), one or more plot points, conflict, characters, etc. A person's SCF (Sentence Similarity Factor) can include gender, role, race, appearance, behavior, speech, expertise, influence, attitude, and emotion. Each SCF includes one or more semantic facts (SFs), which exist as a few words, phrases, or short sentences describing the text content. Because SFs are inherently concise, existing text similarity assessment tools and techniques can be used, such as word embeddings, sentence embeddings, semantic textual similarity (STS), semantic graph matching, and pre-trained language models. STS evaluation is based on TRS (Translation Relationship Standard) matching. Each of these techniques is equally applicable to SCF feature vector similarity evaluation. Depending on the specific application, available resources, and the level of granularity required to measure sentence similarity, one of these text similarity techniques can be selected.
[0034] According to one embodiment, method 100 further includes: obtaining fingerprint recognition instructions that define a fingerprint structure for the text content according to a literature analysis guide; loading the text content into a generative pre-trained transformer (GPT) for literature analysis of the text content; and obtaining semantic facts about the text content by querying the GPT to describe the text content using one or more words according to the fingerprint recognition instructions. Based on the literature analysis of the text content using the GPT, fingerprint recognition instructions are obtained to generate a fingerprint structure for the text content. To generate the fingerprint structure, the text content is loaded into the GPT for literature analysis. By performing literature analysis on the text content, one or more semantic facts can be generated, which describe the text content using one or more words according to the fingerprint recognition instructions.
[0035] At step 104, method 100 further includes: creating a transformation resilient signature (TRS) of the text content, wherein the TRS comprises one or more SCFs of the text content. One or more SCFs of the text content (more specifically, one or more SFs of each SCF of the text content) are used to create the TRS of the text content. The TRS is expected to persist after lexical processing (e.g., trimming, rewriting, and interpreting of the text content). The TRS incorporates several newly introduced semantic projections (i.e., bibliographic analysis elements), also known as SCFs. Adding more SCFs to the TRS makes it more robust. Therefore, hiding the semantics of the text content becomes very difficult.
[0036] According to one embodiment, semantic facts include one or more of the following: key information, plot points, unique facts, and themes within the text content. These semantic facts, such as key information, plot points, unique facts, and themes, are used to create paragraph signatures, also known as transformation resilient signatures (TRS). For example, key information in a typical Cinderella story includes: - Despite numerous obstacles, a good life hope It will always be achieved - perseverance and resilience It is very important for achieving long-term goals. Even towards those who abuse you, maintain your composure. Kindness and tolerance You will also be rewarded. - Inner beauty More important than appearance, etc.
[0037] According to one embodiment, the TRS has a tree structure, wherein the internal nodes of the tree structure include SCFs, and the leaf nodes of the tree structure include semantic facts corresponding to the SCFs. In one implementation, the TRS may have a tree structure with SCFs as internal nodes and SFs as leaf nodes. For example, Figure 4 An exemplary scenario is described in detail.
[0038] According to one embodiment, method 100 further includes: obtaining a nested SCF that includes one or more other SCFs; the TRS tree structure includes internal nodes, which include nested SCFs. Key features such as storylines, character lists, etc. (i.e., SCFs) can include a list of other key features consisting of one or more semantic facts, these other key features are called nested semantic features. Nested SCFs are represented by internal nodes in the TRS tree structure that include one or more other SCFs. For example, Figure 7 An exemplary scenario is described in detail.
[0039] According to one embodiment, the root node of the TRS tree structure includes one or more of the following: the title of the text content, the identifier (ID) of the owner of the text content, and information. Each of the other nodes in the TRS tree structure includes the type, ID, and description of the corresponding SCF or semantic fact of the text content. In the TRS tree structure, the root node represents the TRS of the text content. The root node includes the title of the text content, the ID of the owner of the text content, and detailed information. Other nodes in the TRS tree structure (e.g., internal nodes and leaf nodes) include the type, ID, and description of the corresponding SCF or semantic fact of the text content.
[0040] Therefore, Method 100, by supporting advanced IP protection technologies such as creating a TRS (text fingerprint or semantic structure) for text content, can prevent plagiarism and unauthorized use of text content. Method 100 performs document analysis using GPT to mine one or more SCFs (Search Functions) of the text content. Subsequently, Method 100 generates a TRS (digital signature) using one or more SCFs of the text content, supporting enhanced IP protection. Furthermore, Method 100 concisely summarizes the essence of the text (i.e., semantics), removing textual embellishments and decorations (i.e., noise) from the text content. Additionally, Method 100 improves and unifies text readability to an acceptable level (e.g., 8 years of education), leading interpreters to use a basic dictionary. Therefore, after testing the transformation, different texts describing the same or similar ideas become identical or very similar. In one implementation scenario, Method 100 can also be deployed in a copyright cloud service. In another implementation scenario, Method 100 can also be deployed as part of a content distribution network (CDN). Method 100 can be used independently or as an enhancement to existing text similarity detection tools and technologies.
[0041] Steps 102 and 104 are merely illustrative. Other alternatives may be provided without departing from the scope of the claims herein, such as adding one or more steps, deleting one or more steps, or providing one or more steps in a different order.
[0042] In one aspect, the present invention provides a computer program product including instructions. When executed by a computer, the instructions cause the computer to perform the semantic content fingerprinting method 100 for text content described above. In another aspect, the present invention provides a computer-readable storage medium including instructions. When executed by a computer, the instructions cause the computer to perform the semantic content fingerprinting method 100 for text content described above.
[0043] Figure 2 This is a block diagram of a semantic content fingerprint recognition device provided in one embodiment of the present invention. Figure 2 Combination Figure 1 The elements within are described. (See reference.) Figure 2 A block diagram 200 of device 202 is shown, which includes a fingerprint recognition engine 204, a generative pre-trained transformer (GPT) 206, a fingerprint combiner 208, a processor 210, and a memory 212. Text content 214 is also shown.
[0044] Device 202 may include suitable logic, circuitry, interfaces, and / or code for performing semantic content fingerprinting. Examples of device 202 may include, but are not limited to, computers, laptops, custom hardware, or any other portable or non-portable electronic devices. Device 202 may also be referred to as a semantic content fingerprinting system.
[0045] GPT-206 can also be called a pre-trained language model. GPT-206 is trained on a large amount of text data, capturing rich semantic information, thus enabling it to effectively measure sentence semantic similarity. GPT-206 can be used with various machine learning (ML) algorithms, such as supervised ML algorithms, unsupervised ML algorithms, deep learning (DL) algorithms, and artificial neural network (ANN) algorithms.
[0046] Processor 210 may include suitable logic, circuitry, interfaces, and / or code for executing instructions stored in memory 212. Examples of processor 210 may include, but are not limited to, integrated circuits, coprocessors, microprocessors, microcontrollers, complex instruction set computing (CISC) processors, application-specific integrated circuit (ASIC) processors, reduced instruction set (RISC) processors, very long instruction word (VLIW) processors, central processing units (CPUs), state machines, data processing units, and other processors or circuits. Furthermore, processor 210 may refer to one or more separate processors, processing devices, or processing units as part of a machine.
[0047] Memory 212 may include suitable logic, circuitry, interfaces, and / or code for storing machine code and / or instructions executable by processor 210. Examples of implementations of memory 212 include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital card (SD), solid-state drive (SSD), computer-readable storage media, and / or CPU cache memory. Memory 212 may store an operating system and / or computer program product to operate device 202. Computer-readable storage media used to provide non-transitory memory may include, but are not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing.
[0048] In operation, device 202 includes a fingerprinting engine 204, which performs literal analysis on text content 214 using a generative pre-trained transformer (GPT) 206 to obtain one or more semantic content features (SCFs) of the text content 214. Each SCF includes one or more semantic facts about the text content 214, and each semantic fact includes one or more words describing the text content 214. GPT 206 performs literal analysis on the text content 214 and generates one or more SCFs of the text content 214. The fingerprinting engine 204 is used to obtain the generated one or more SCFs of the text content 214. For example, Figure 1 Examples of one or more SCFs are detailed below. Each SCF includes one or more semantic facts, which are presented in the form of short sentences describing the text content 214.
[0049] According to one embodiment, the fingerprinting engine 204 is used to: obtain fingerprinting instructions that define a fingerprint structure for text content 214 according to a literature analysis guide; load the text content 214 into a generative pre-trained transformer (GPT) 206 for literature analysis of the text content 214; and obtain semantic facts about the text content 214 by querying the GPT 206 to describe the text content 214 using one or more words according to the fingerprinting instructions. By loading the text content 214 (i.e., the original input) into the GPT 206, the fingerprinting engine 204 is responsible for managing the process of creating one or more SCFs. The GPT 206 is used to create a query sequence to create a comprehensive set of SCFs and SFs by performing literature analysis on the text content 214. Furthermore, the fingerprint structure describes the instructions that the fingerprinting engine 204 executes to create the fingerprint. In one implementation scenario, the fingerprint structure can reflect a multi-level tree similar to an SCF organization, where the leaf nodes describe the GPT 206 queries required to mine appropriate facts.
[0050] The following example illustrates the definition of a fingerprint structure: fingerprint sequence Key information: What are the key information in this paragraph? Genre: Fairy Tale Characters: Who are the main characters? figure: What are the unique characteristics of this character? What similarities does he share with other figures? Plot: What is the plot? event: Can you find this event in other stories? ... Apparatus 202 also includes a fingerprint combiner 208 for creating a transformation resilient signature (TRS) of text content 214, wherein the TRS includes one or more SCFs of text content 214. The fingerprint combiner 208 combines one or more SCFs of text content 214 created by GPT 206. The one or more SCFs represent text or alternative text representations (ontologies) that support text similarity evaluation. The TRS (i.e., fingerprint) of text content 214 aggregates one or more SCFs, each SCF including one or more semantic facts of text content 214.
[0051] According to one embodiment, semantic facts include one or more of the following: key information in the text content, plot points, unique facts, and topics. For example, Figure 1The examples detail one or more key pieces of information.
[0052] According to one embodiment, fingerprint combiner 208 is used to create a TRS with a tree structure, wherein the internal nodes of the tree structure include SCFs, and the leaf nodes of the tree structure include semantic facts corresponding to the SCFs. For example, Figure 4 An exemplary scenario is described in detail.
[0053] According to one embodiment, fingerprint recognition engine 204 is used to acquire nested SCFs that include one or more other SCFs; fingerprint combiner 208 is used to create internal nodes of a TRS tree structure, wherein the internal nodes include nested SCFs. Facts that can be described by other facts (e.g., key people, timeline events, etc.) are represented by key features. Such key features that include these facts are called nested semantic features (NSFs). For example, Figure 7 An exemplary scenario is described in detail.
[0054] According to one embodiment, the root node of the TRS tree structure includes one or more of the following: the title of the text content, the identifier (ID) of the owner of the text content, and information. Each of the other nodes in the TRS tree structure includes the type, ID, and description of the corresponding SCF or semantic fact of the text content. In the TRS tree structure, the root node represents the TRS of text content 214. The root node includes the title of text content 214, the ID of the owner of text content 214, and detailed information. The other nodes of the TRS tree structure (e.g., internal nodes and leaf nodes) include the type, ID, and description of the corresponding SCF or semantic fact of text content 214.
[0055] Figure 3 This is an illustration of a semantic content fingerprinting system provided in an embodiment of the present invention. Figure 3 Combination Figure 1 and Figure 2 The elements within are described. (See reference.) Figure 3 A block diagram 300 of a semantic content fingerprinting system 302 is shown, which includes a GPT 304 and a fingerprinting engine 306. The fingerprinting engine 306 includes a combiner 308. A fingerprint structure 310, a TRS 312, and text content 314 are also shown. The TRS 312 is indicated by a dashed box, which is for illustrative purposes only.
[0056] The semantic content fingerprinting system 302, GPT 304, fingerprint recognition engine 306, combiner 308, and text content 214 correspond to device 202, GPT 206, fingerprint recognition engine 204, fingerprint combiner 208, and text content 314, respectively. TRS 312 can also be referred to as the fingerprint of text content 314.
[0057] The fingerprinting engine 306 is used to load text content 314 (also referred to as input) into GPT 304 and provide GPT 304 with multiple queries (or execution instructions) about the text content 314, with the aim of retrieving semantic features from GPT 304. GPT 304 is used to create one or more SCFs of text content 314 through document analysis. GPT 304 can also be used to use a large language model (LLM) when performing document analysis on text content 314. Subsequently, one or more SCFs of text content 314 (also referred to as output) are provided to the fingerprinting engine 306, where each SCF includes one or more semantic facts about text content 314. For example, Figure 1 and Figure 2 These semantic facts have been detailed. The fingerprinting engine 306 is also used to obtain fingerprinting instructions defining a fingerprint structure 310 for the text content 314 according to the document analysis guidelines. The fingerprint structure 310 is constructed recursively according to the response iteration of GPT 304. Regular and combined SCFs are processed appropriately. Subsequently, the combiner 308 (i.e., the fingerprint combiner) is used to create a TRS 312 for the text content 314. The TRS 312 adopts a tree structure. For example, Figure 4 This tree structure is described in detail.
[0058] Figure 4 The tree structure of a Transformation Resilient Signature (TRS) for text content provided in one embodiment of the present invention is shown. Figure 4 Combination Figure 1 , Figure 2 and Figure 3 The elements within are described. (See reference.) Figure 4 The tree structure 400 of the TRS for the text content is shown (e.g., Figure 2(Text content 214 in the text). The tree structure 400 includes a root node 402 and multiple internal nodes, such as the first internal node 404A, the second internal node 404B, up to the Nth internal node 404N. Each of the multiple internal nodes includes multiple leaf nodes. For example, the first internal node 404A includes the first multiple leaf node 406A, the second internal node 404B includes the second multiple leaf node 406B, and the Nth internal node 404N includes the Nth multiple leaf node 406N.
[0059] The root node 402 of the tree structure 400 represents the TRS of the text content, and each of the multiple internal nodes represents the SCF of the text content. For example, the first internal node 404A represents the first SCF of the text content (i.e., SCF1), the second internal node 404B represents the second SCF of the text content (i.e., SCF2), and the Nth internal node 404N represents the nth SCF of the text content (i.e., SCFn). Furthermore, the first multiple leaf node 406A associated with the first internal node 404A represents one or more semantic facts (SFs) of the text content. The first multiple leaf node 406A can include 1 to m leaf nodes, thus representing m SFs of the text content (i.e., SFm). Similar to the first multiple leaf node 406A, the second multiple leaf node 406B and the Nth multiple leaf node 406N represent m SFs of the text content.
[0060] Figure 5 This is a tree structure of text content TRS provided in another embodiment of the present invention. Figure 5 Combination Figure 1 , Figure 2 , Figure 3 and Figure 4 The elements within are described. (See reference.) Figure 5 The diagram illustrates the tree structure 500 of the TRS (Text Representation System) for the text content. The tree structure 500 includes a root node 502 and multiple internal nodes, such as a first internal node 504A, a second internal node 504B, and so on up to an Nth internal node 504N. Each of the multiple internal nodes includes multiple leaf nodes. For example, the first internal node 504A includes a first multiple leaf node 506A, the second internal node 504B includes a second multiple leaf node 506B, and the Nth internal node 504N includes an Nth multiple leaf node 506N.
[0061] The root node 502 of the tree structure 500 represents the TRS (fingerprint) of the text content, and each of the multiple internal nodes represents the SCF (Signal Credential Factor) of the text content. For example, the first internal node 504A represents a "fork" in the text content, the second internal node 504B represents a "loop" in the text content, and the Nth internal node 504N represents an "end" in the text content. Furthermore, the first multiple leaf node 506A associated with the first internal node 504A represents one or more details (i.e., minor or secondary details) of the text content. Similarly, the second multiple leaf node 506B and the Nth multiple leaf node 506N represent one or more details of the text content.
[0062] Figure 6 This is an exemplary scenario of a tree structure for text content TRS provided in another embodiment of the present invention. Figure 6 Combination Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 The elements within are described. (See reference.) Figure 6 The diagram illustrates a tree structure 600 of the TRS text content. The tree structure 600 includes a root node 602 and multiple internal nodes, such as a first internal node 604A, a second internal node 604B, and a third internal node 604C. Each of the multiple internal nodes includes multiple leaf nodes. For example, the first internal node 604A includes a first multiple leaf node 606A, the second internal node 604B includes a second multiple leaf node 606B, and the third internal node 604C includes a third multiple leaf node 606C.
[0063] The root node 602 of the tree structure 600 represents the TRS, i.e., "Cinderella" in the fairy tale. Each of the multiple internal nodes represents a SCF (Site-Fact) of the fairy tale. For example, the first internal node 604A represents the "plot" (i.e., SCF1) of the fairy tale, the second internal node 604B represents the "key information" (i.e., SCF2) of the fairy tale, and the third internal node 604C represents the "unique fact" (i.e., SCF3) of the fairy tale. Furthermore, the first multiple leaf node 606A associated with the first internal node 604A represents one or more SFs (e.g., details about "the death of the father" and "the ball"). Similar to the first multiple leaf node 606A, the second multiple leaf node 606B and the third multiple leaf node 606C represent other minor details of the fairy tale. For example, the second multiple leaf node 606B represents an SF associated with "key information" (e.g., hope and perseverance), and the third multiple leaf node 606C represents an SF associated with "unique facts" (e.g., the glass slipper and the shoe-fitting ID).
[0064] Figure 7 This is a tree structure of text content TRS provided in another embodiment of the present invention. Figure 7 Combination Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 The elements within are described. (See reference.) Figure 7 The diagram illustrates the tree structure 700 of the TRS text content. The tree structure 700 includes a root node 702 and multiple internal nodes, such as the i-th internal node 704I, the j-th internal node 704J, and so on up to the n-th internal node 704N. The j-th internal node 704J includes internal node 704J1. Each of the multiple internal nodes includes multiple leaf nodes. For example, the i-th internal node 704I includes the first multiple leaf node 706A, and internal node 704J1 includes the second multiple leaf node 708A. Similarly, the n-th internal node 704N may include the n-th multiple leaf node (not shown here).
[0065] The root node 702 of the tree structure 700 represents the TRS of the text content, and each of the multiple internal nodes represents the SCF of the text content. For example, the i-th internal node 704I represents the i-th SCF (i.e., SCFi) of the text content, the j-th internal node 704J represents the j-th SCF (i.e., SCFj) of the text content, and the n-th internal node 704N represents the n-th SCF (i.e., SCFn) of the text content. Since the j-th internal node 704J includes internal node 704J1, the j-th internal node 704J is called a nested internal node, representing nested SCFs of the text content. Furthermore, the first plurality of leaf nodes 706A associated with the i-th internal node 704I represents one or more semantic facts (SFs) of the text content. The first plurality of leaf nodes 706A can include 1 to m leaf nodes, and therefore can represent m SFs (i.e., SFm) of the text content. Similar to the first multiple leaf node 706A, the second multiple leaf node 708A associated with the internal node 704J1 represents m SFs (i.e., SFm) of the text content.
[0066] Figure 8 This is a flowchart of semantic content fingerprint recognition provided in one embodiment of the present invention. Figure 8 Combination Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 The elements within are described. (See reference.) Figure 8 A flowchart 800 is shown, including operations 802 to 816.
[0067] Initially, the keyword "start" indicated the start of flowchart 800.
[0068] At operation 802, information is sent to the fingerprint recognition engine (e.g., Figure 3 The fingerprint recognition engine 306 in the system provides an input sample, which forwards the input sample to GPT (e.g., Figure 3 GPT 304 in (the document).
[0069] At operation 804, the fingerprinting engine provides multiple instructions to the GPT regarding the input sample, with the aim of extracting semantic content features from the GPT.
[0070] At operation 806, GPT is used to provide one or more SCFs of input samples to the fingerprint recognition engine.
[0071] At operation 808, one or more SCFs provided are added as one or more internal nodes to the tree structure of the input sample's TRS.
[0072] At operation 810, check for any nested SCFs. If one or more nested SCFs are found, repeat operations 804 through 810.
[0073] At operation 812, the fingerprint engine provides multiple instructions to the GPT with the aim of obtaining semantic facts (SF) from the GPT.
[0074] At operation 814, GPT is used to provide one or more SFs of input samples to the fingerprint recognition engine.
[0075] At operation 816, one or more SFs provided are added to the tree structure of the input sample's TRS as one or more leaf nodes or one or more internal nodes.
[0076] The keyword "end" indicates the end of flowchart 800.
[0077] Modifications to the embodiments of the invention described above may be made without departing from the scope of the invention as defined by the appended claims. The terms “comprising,” “incorporated,” “having,” “is,” and the like used to describe and claim the invention are intended to be interpreted in a non-exclusive manner, meaning that items, components, or elements not explicitly described may also be present. Singular references should also be interpreted as relating to the plural. The word “exemplary” as used herein means “as an example, instance, or illustration.” Any embodiment described as “exemplary” is not necessarily to be construed as being more preferred or advantageous than other embodiments, and / or excluding combinations of features from other embodiments. The word “optionally” as used herein means “provided in some embodiments and not in others.” It should be understood that certain features of the invention described in the context of a single embodiment for clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for brevity may also be provided individually or by any suitable combination or, where appropriate, in any other described embodiment of the invention.
Claims
1. A semantic content fingerprinting method (100), characterized in that, The method (100) includes: By using a generative pre-trained transformer (GPT) (206) to perform literature analysis on the text content (214), one or more semantic content features (SCFs) of the text content (214) are obtained, wherein each SCF includes one or more semantic facts about the text content (214), and each semantic fact includes one or more words describing the text content (214); Create a transformation resilient signature (TRS) (312) for the text content (214), wherein the TRS (312) includes one or more SCFs of the text content (214).
2. The method (100) according to claim 1, characterized in that, Also includes: According to the literature analysis guide, obtain the fingerprint recognition instructions that define the fingerprint structure (310) for the text content (214); The text content (214) is loaded into the generative pre-trained transformer (GPT) (206) used for literature analysis of the text content (214). By querying the GPT (206) to describe the text content (214) using one or more words according to the fingerprinting instructions, semantic facts about the text content (214) are obtained.
3. The method (100) according to claim 1 or 2, characterized in that, The semantic facts include one or more of the key information, plot points, unique facts, and themes in the text content (214).
4. The method (100) according to any one of claims 1 to 3, characterized in that, The TRS (312) has a tree structure, wherein the internal nodes of the tree structure include the SCF, and the leaf nodes of the tree structure include the semantic facts corresponding to the SCF.
5. The method (100) according to claim 4, characterized in that, Also includes: Retrieves a nested SCF that includes one or more other SCFs; The TRS tree structure includes internal nodes, and the internal nodes include the nested SCF.
6. The method (100) according to claim 4 or 5, characterized in that, The root node (402) of the TRS tree structure includes one or more of the title of the text content (214), the identifier (ID) of the owner of the text content (214), and information, while each of the other nodes in the TRS (312) tree structure includes the type, ID, and description of the corresponding SCF or semantic fact of the text content (214).
7. A semantic content fingerprint recognition device (202), characterized in that, The device (202) includes: The fingerprint recognition engine (204) is used to perform literature analysis on the text content (214) by using a generative pre-trained transformer (GPT) (206) to obtain one or more semantic content features (SCFs) of the text content (214), wherein each SCF includes one or more semantic facts about the text content (214), and each semantic fact includes one or more words describing the text content (214); A fingerprint combiner (208) is used to create a transformation resilient signature (TRS) (312) of the text content (214), wherein the TRS (312) includes one or more SCFs of the text content (214).
8. The apparatus (202) according to claim 7, characterized in that, The fingerprint recognition engine (204) is used for: According to the literature analysis guide, obtain the fingerprint recognition instructions that define the fingerprint structure (310) for the text content (214); The text content (214) is loaded into the generative pre-trained transformer (GPT) (206) used for literature analysis of the text content (214). By querying the GPT (206) to describe the text content (214) using one or more words according to the fingerprinting instructions, semantic facts about the text content (214) are obtained.
9. The apparatus (202) according to claim 7 or 8, characterized in that, The semantic facts include one or more of the key information, plot points, unique facts, and themes in the text content (214).
10. The apparatus (202) according to any one of claims 7 to 9, characterized in that, The fingerprint combiner (208) is used to create the TRS (312) with a tree structure, wherein the internal nodes of the tree structure include the SCF, and the leaf nodes of the tree structure include the semantic facts corresponding to the SCF.
11. The apparatus (202) according to claim 10, characterized in that, The fingerprint recognition engine (204) is used to acquire a nested SCF that includes one or more other SCFs; the fingerprint combiner (208) is used to create internal nodes of the TRS tree structure, wherein the internal nodes include the nested SCFs.
12. The apparatus (202) according to claim 10 or 11, characterized in that, The root node (402) of the TRS tree structure includes one or more of the title of the text content (214), the identifier (ID) of the owner of the text content (214), and information, while each of the other nodes of the TRS tree structure includes the type, ID, and description of the corresponding SCF or semantic fact of the text content (214).
13. A computer program product comprising instructions, characterized in that, When the computer program product is executed by a computer, the instructions cause the computer to perform the semantic content fingerprinting method (100) for text content (214) according to any one of claims 1 to 6.
14. A computer-readable storage medium including instructions, characterized in that, When the instructions are executed by a computer, the instructions cause the computer to perform the semantic content fingerprinting method (100) for text content (214) according to any one of claims 1 to 6.