AI generation content front-end real-time identification method and system based on bimodal verification
By obtaining file information and operation data in the front-end environment of the user equipment and performing dual-modal verification, the AI generated content recognition method solves the problems of single detection dimension, high computing cost, high privacy risks and poor real-time performance, and achieves efficient, accurate, secure and real-time AI generated content recognition.
Patent Information
- Application Number
- CN202510625682.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art has problems such as single detection dimension, high computing cost, high privacy risks and poor real-time performance when detecting content generated by AI, which is difficult to meet the needs of high precision, low resource consumption, strong privacy protection and fast response.
The AI-generated content front-end real-time recognition method is adopted based on dual-modal verification. User file information and operation data are obtained in the front-end environment of the user equipment, behavior fingerprints are generated and edge density and word frequency jump detection are performed. Combined with the preset dual-modal risk determination engine, risk scores are evaluated and locally identified.
It improves detection accuracy, reduces computing resource consumption, avoids the risk of data leakage, realizes fast real-time interception, improves user experience, and is suitable for application scenarios with high real-time requirements such as social media platforms.
Smart Images

Figure CN120495857A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of data security technology, and specifically to a method and system for front-end real-time recognition of AI-generated content based on bimodal verification. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, AI-generated content (such as deepfakes and machine-generated text) is becoming increasingly widespread online. As the technology used to generate this content continues to advance, it becomes increasingly difficult to distinguish it from real content, posing significant challenges to verifying its authenticity. For example, on social media platforms, fake AI-generated images or text can mislead users and even be used for malicious purposes, such as spreading disinformation and online fraud. Currently, technologies for detecting AI-generated content fall into two main categories: content feature analysis and user authentication. Content feature analysis relies on training AI models (such as convolutional neural networks and natural language processing models) to analyze the statistical characteristics of the file itself. However, this technology relies solely on file features and cannot distinguish between "quick human input" and "machine-generated content," making it easily circumvented by carefully crafted AI-generated content. Furthermore, content analysis typically requires the deployment of large AI models, which are computationally expensive and difficult to run in front-end environments such as browsers. User identity verification technology, meanwhile, indirectly assesses user credibility through methods like account reputation ratings, mobile phone verification, and biometric authentication, thereby inferring the authenticity of content. However, this approach also has significant flaws: identity verification technology can be easily circumvented by forged identities. For example, unauthorized individuals can purchase real-name accounts to publish false content. Furthermore, existing detection technologies pose privacy risks. Cloud-based detection requires uploading user files, posing a risk of data leakage, such as uploading ID images to third-party servers. Secondly, cloud-based detection lacks real-time performance, with latency affecting the user experience. For example, image uploads on social media platforms may experience lag.
[0003] Existing technologies for detecting AI-generated content suffer from problems such as a single detection dimension, high computational cost, significant privacy risks, and poor real-time performance. These issues make it difficult to meet the demands of practical applications for high accuracy, low resource consumption, strong privacy protection, and rapid response. Therefore, a front-end real-time recognition method for AI-generated content that can improve recognition accuracy is urgently needed. Summary of the Invention
[0004] The embodiments of this specification provide a method and system for front-end real-time recognition of AI-generated content based on dual-modal verification, and the technical solution is as follows: On the first aspect, the embodiments of this specification provide a front-end real-time identification method for AI-generated content based on bimodal verification, including: based on the front-end environment running on the user device, obtaining user file information and user operation data, and generating a behavioral fingerprint based on the user operation data; performing edge density processing and word frequency jump detection processing on the user file information to obtain content features of the user file information; based on a preset bimodal risk determination engine, determining the user's risk score data based on the content features and user operation data; when the risk score data meets the first preset condition, determining the user file information as non-AI generated content, and adding the behavioral fingerprint to the metadata corresponding to the user file information to obtain behavioral fingerprint metadata.
[0005] On the second aspect, an embodiment of this specification provides a front-end real-time identification system for AI-generated content based on bimodal verification, including: a data acquisition module, which is used to obtain user file information and user operation data based on the front-end environment running on the user device, and generate a behavioral fingerprint based on the user operation data; a feature extraction module, which is used to perform edge density processing and word frequency jump detection processing on the user file information to obtain the content characteristics of the user file information; a risk judgment module, which is used to determine the user's risk score data based on the content characteristics and user operation data based on a preset bimodal risk judgment engine; a metadata marking module, which is used to determine the user file information as non-AI generated content when the risk score data meets the first preset condition, and add the behavioral fingerprint to the metadata corresponding to the user file information to obtain behavioral fingerprint metadata.
[0006] The beneficial effects of the technical solutions provided by some embodiments of this specification include at least: The embodiments of this specification adopt a dual-modal verification method of user behavior and content, comprehensively considering user operating habits and file content characteristics, thereby improving detection accuracy. It is difficult for machine-generated content to simultaneously forge human operating habits and natural content characteristics. Therefore, the embodiments of this specification can more effectively identify AI-generated content through dual-modal verification and reduce the misjudgment rate.
[0007] Moreover, the embodiments of this specification run on the front end without relying on large AI models, and achieve lightweight detection through a rule engine, which greatly reduces computing resource consumption and avoids network delays and data transmission costs caused by cloud computing.
[0008] Furthermore, the detection latency of the embodiments of this specification is reduced from 1200ms in cloud-based detection to 80ms, and the response time is less than 100ms, enabling rapid, real-time interception. This eliminates network transmission and cloud-based queuing time, allowing users to upload images or text without noticeable lag, significantly improving the user experience. This is particularly suitable for applications requiring high real-time performance, such as content publishing on social media platforms.
[0009] Furthermore, the embodiments of this specification can be run locally, with user behavior and file content processed locally without uploading to the cloud, effectively avoiding the risk of data leakage. By relying on the browser's native API for operation, it complies with privacy regulations such as GDPR, fully protecting user privacy and data security.
[0010] In addition, the embodiments of this specification comprehensively analyze user behavior and content characteristics, so that attackers need to simultaneously forge content characteristics and simulate human operation behavior, greatly increasing the cost of forgery. Uploaders of AI-generated content typically present the characteristics of "high content quality + low operation complexity", while real users often have multiple adjustments and random operation paths. The design of the present invention can effectively identify this difference, thereby enhancing the ability to identify AI-generated content and reducing the risk of being bypassed.
[0011] The embodiments of this specification use innovative dual-modal verification technology to solve the problems of single detection dimension, high computational cost, high privacy risk and poor real-time performance in the existing technology, and provide an efficient, accurate, secure and real-time solution for the recognition of AI-generated content, which has broad application prospects and important practical significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 This is a schematic diagram of an application scenario of a front-end real-time recognition method for AI-generated content based on bimodal verification provided in this manual.
[0014] Figure 2 This is a flowchart of a method for real-time front-end recognition of AI-generated content based on bimodal verification provided in this manual.
[0015] Figure 3 This is a flowchart of generating behavioral fingerprints based on user operation data provided in this manual.
[0016] Figure 4 This is a flowchart of obtaining the content characteristics of user file information provided in this manual.
[0017] Figure 5 This is a structural diagram of a front-end real-time recognition system for AI-generated content based on dual-modal verification provided in this manual. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of this specification will be described clearly and completely below in conjunction with the drawings in the embodiments of this specification.
[0019] Throughout this specification, the claims, and the accompanying drawings, the terms "first," "second," and so forth are used to distinguish between different items, not to describe a particular order. Furthermore, the term "comprises" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may include other steps or elements inherent to the process, method, product, or apparatus.
[0020] The multiple embodiments of this specification provide a method for real-time identification of AI-generated content front-end based on dual-modal verification. The executor of the method can be the system for real-time identification of AI-generated content front-end based on dual-modal verification provided in an embodiment of the present invention.
[0021] Before this specification elaborates on the method for real-time front-end recognition of AI-generated content based on dual-modal verification in combination with one or more embodiments, it first introduces the application scenario of the method for real-time front-end recognition of AI-generated content based on dual-modal verification.
[0022] See also Figure 1 , Figure 1 A schematic diagram of an application scenario for the method for real-time front-end recognition of AI-generated content based on bimodal verification provided by an embodiment of the present invention. In this embodiment, the system 100 for real-time front-end recognition of AI-generated content based on bimodal verification may include a terminal 110, which may be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC). Terminal 110 is provided with a front-end environment, that is, terminal 110 has a client application, browser, or other front-end framework that can run on the user's device. For example, a browser such as Chrome, Firefox, Safari, etc., through which users access web applications; for example, a client application such as a mobile application (iOS, Android) or a desktop application (Windows, macOS); for example, other front-end frameworks such as a desktop application based on Electron or a cross-platform application based on React Native.
[0023] In the present embodiment, terminal 110 may be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC). Terminal 110 includes a central processing unit (CPU), a graphics processing unit (GPU), memory, storage device, network communication module, sensor, display screen, battery and power management module, etc. The central processing unit can execute logical calculations, resource scheduling and rendering optimization algorithms of front-end applications; the graphics processor can be used to accelerate front-end page rendering, especially the rendering of complex graphics, animations and multimedia content; the memory can be used to store the runtime data, resource files, user behavior data and environmental perception data of front-end applications; the storage device can be used to store the code, resource files and cached data of front-end applications; the network communication module can be used to communicate with the server to obtain dynamic resources and real-time environmental information; sensors can be used to perceive the environmental status of user devices in real time, such as network sensors for obtaining network bandwidth, latency and connection status, etc., and performance sensors for obtaining the CPU, GPU and memory usage of terminal 110, etc., and screen sensors for obtaining screen resolution, brightness and refresh rate, etc.; the display screen can be used to present the content of front-end applications; the battery and power management module can provide power support for device operation and optimize energy consumption.
[0024] The terminal 110 of the embodiment of this description can obtain user file information and user operation data based on the front-end environment running on the user device, and generate a behavioral fingerprint based on the user operation data; perform edge density processing and word frequency jump detection processing on the user file information to obtain content characteristics of the user file information; based on a preset bimodal risk judgment engine, determine the user's risk score data according to the content characteristics and user operation data; when the risk score data meets the first preset condition, the user file information is judged as non-AI generated content, and the behavioral fingerprint is added to the metadata corresponding to the user file information to obtain behavioral fingerprint metadata.
[0025] It should be noted that Figure 1 The scenario diagram of the AI-generated content front-end real-time recognition system 100 based on dual-modal verification is only an example. The AI-generated content front-end real-time recognition system based on dual-modal verification and the scenario described in the embodiment of the present invention are intended to more clearly illustrate the technical solution of the embodiment of the present invention, and do not constitute a limitation on the technical solution provided by the embodiment of the present invention. Ordinary technicians in this field can know that with the evolution of the AI-generated content front-end real-time recognition system based on dual-modal verification and the emergence of new scenarios, the technical solution provided by the embodiment of the present invention is also applicable to similar technical problems.
[0026] See also Figure 2 , Figure 2 This is a flow chart of a method for real-time identification of AI-generated content front-end based on dual-modal verification provided by an embodiment of the present invention. The method for real-time identification of AI-generated content front-end based on dual-modal verification can be performed by Figure 1 The AI-generated content front-end real-time identification system 100 based on dual-modal verification is executed. The AI-generated content front-end real-time identification method based on dual-modal verification can at least include the following steps: 200. Based on the front-end environment running on the user device, obtain user file information and user operation data, and generate a behavior fingerprint based on the user operation data.
[0027] In this embodiment, terminal 110 is provided with a front-end environment, such as a browser (e.g., Chrome, Firefox, Safari, etc.) through which users access web applications; client applications (e.g., mobile applications (iOS, Android) or desktop applications (Windows, macOS); and other front-end frameworks (e.g., desktop applications based on Electron and cross-platform applications based on React Native). User file information may include images, text, and other information. User operation data refers to the user's behavioral data on user file information within the terminal's front-end environment. The behavioral fingerprint may be a hash value generated based on the user operation data.
[0028] In some embodiments, the user operation data may include the coordinates and timing of the mouse movement path, the number of file edits, and the operation time.
[0029] In some embodiments, see Figure 3 , Figure 3 The following is a flow chart of generating a behavior fingerprint based on user operation data according to an embodiment of the present invention. Generating a behavior fingerprint based on user operation data includes: 2000. Obtain a random projection matrix, where any element in the random projection matrix is randomly sampled from a standard normal distribution; 2010, Determining behavioral data vectors through user operation data; 2020. Multiply the behavior data vector by the random projection matrix to obtain a low-dimensional projection vector; 2030. Determine the hash value of any dimensional element in the low-dimensional projection vector based on the hash determination engine; 2040. Generate a behavioral fingerprint based on the hash value of any dimensional element in the low-dimensional projection vector.
[0030] In this embodiment, when the terminal 110 generates a behavioral fingerprint based on user operation data, it can first obtain a d×n dimensional random projection matrix, where any element in the random projection matrix is obtained by random sampling from a standard normal distribution; then, the behavioral data vector is determined based on the user operation data, and the behavioral data vector is multiplied by the random projection matrix to obtain a low-dimensional projection vector; for the low-dimensional projection vector, the hash value can be calculated bit by bit based on the hash determination engine to calculate the behavioral fingerprint.
[0031] In some embodiments, based on a hash determination engine, the hash value of any dimensional element in a low-dimensional projection vector is determined, including: when the dimensional element is less than a first preset value, the hash value corresponding to the any dimensional element is a first value; when the dimensional element is not less than the first preset value, the hash value corresponding to the any dimensional element is a second value.
[0032] For example, for a low-dimensional projection vector p, calculate any dimension element p i The hash value of dimension element p i When it is less than 0, the dimension element p i The hash value of the dimension element p is set to 0; i When not less than 0, the dimension element p i The hash value of is set to 1. For example, for the low-dimensional projection vector p=[0.5,-0.3,1.2,...,-0.8], the hash value corresponding to the low-dimensional projection vector p is [1,0,1,...,0].
[0033] The embodiment of this specification can also convert the hash value corresponding to the low-dimensional projection vector p into a hexadecimal string, for example: A3F9C7D2. The embodiment of this specification can quickly compare the similarity of user behaviors by generating a behavioral fingerprint (such as a 16-bit string).
[0034] 210. Perform edge density processing and word frequency jump detection processing on the user file information to obtain content features of the user file information.
[0035] In some embodiments, see Figure 4 , Figure 4 This is a flow chart of obtaining content features of user file information provided by an embodiment of the present invention. Content features include image features and text features. Edge density processing and word frequency jump detection processing are performed on the user file information to obtain content features of the user file information, including: 2100. Obtain image information and text information of user file information; 2110. Perform size processing on the image in the image information to obtain an image with a preset size; 2120. Calculate the edge density of an image having a preset size, where the edge density is an image feature corresponding to the user file information; 2130. Obtain the abnormal punctuation rate corresponding to the text information; 2140. Perform word frequency jump detection processing on the text information to obtain word frequency jump detection data; 2150. Determine a context coherence score based on word frequency jump detection data; the abnormal punctuation rate and the context coherence score are text features corresponding to the user file information.
[0036] In this embodiment, the terminal 110 may obtain an image with a preset size by performing size processing on the image in the image information. For example, the preset size may be set to 100×100 pixels.
[0037] In some embodiments, calculating the edge density of an image with a preset size includes: obtaining the number of edge pixels and the total number of pixels of the image with the preset size, and the edge density is the ratio of the number of edge pixels to the total number of pixels.
[0038] In some embodiments, word frequency jump detection processing is performed on text information to obtain word frequency jump detection data, including: based on a preset sliding window, sliding the sliding window to process the text information, and determining the number of occurrences of each word in each sliding window; determining the word frequency change of each word in an adjacent sliding window based on the number of occurrences of each word; and obtaining word frequency jump detection data through the word frequency change.
[0039] In this embodiment, word frequency jump detection is performed on text information. This involves determining the frequency of each word in a sliding window and observing how these frequencies vary across windows. The specific steps include: counting word frequencies (i.e., counting the number of times each word appears in each sliding window); and calculating jump values (i.e., comparing the frequency change of the same word in adjacent sliding windows). If a word's frequency changes significantly across adjacent windows (i.e., "jumps"), the word's occurrence pattern is considered likely to be unnatural text.
[0040] For example, assuming the sliding window size is 3 and the text is "This is a test test text", the results of word frequency jump detection include the first sliding window: "This is a test", the word frequency is: {"this": 1, "is": 1, "one": 1, "test": 1}; the second sliding window: "a test test", the word frequency is: {"one": 1, "test": 2}; the third sliding window: "test test text", the word frequency is: {"test": 2, "text": 1}.
[0041] In this example, the word "test" increases in frequency from 1 to 2 in the second window, and then remains at 2 in the third window. This rapid change in frequency may indicate anomalies in the text, such as excessive repetition or a lack of natural contextual transitions.
[0042] The embodiments of this specification can quantitatively score the coherence of a text based on the results of word frequency jump detection. The specific calculation method for the score can be: defining a jump threshold, setting a threshold for word frequency jumps (for example, a word frequency change exceeding 10%); then counting the number of jumps, that is, counting the number of word frequency jumps exceeding the threshold in all sliding windows in the text; and then calculating the coherence score, that is, calculating the coherence score based on the jump threshold and the number of jumps. A greater number of jumps results in a lower coherence score; a lower number of jumps results in a higher coherence score.
[0043] For example, the embodiment of this specification can set the jump threshold to 10%. In the text "This is a test test text", the word frequency change rate of "test" from the first window to the second window is 100% (from 1 to 2), which exceeds the threshold of 10%. The word frequency change rate of "test" from the second window to the third window is 0% (remains unchanged). The word frequency change of "text" from the second window to the third window is 1. Although the change rate cannot be calculated, it can be regarded as a significant change. Therefore, the word frequency jump of the word "test" is obvious, which may indicate that there is an abnormality in the text. That is, the jump in the word frequency of "test" exceeds the threshold, so the coherence score is low. However, the word frequency change of natural text (such as "This is a test text used to detect context coherence") is relatively gentle, and the coherence score is higher.
[0044] 220. Based on the preset bimodal risk determination engine, the user's risk score data is determined according to the content characteristics and user operation data.
[0045] In some embodiments, based on a preset bimodal risk determination engine, the user's risk score data is determined according to content features and user operation data, including: based on the content anomaly determination rules in the bimodal risk determination engine, determining the content risk score data according to the content features; based on the behavior anomaly determination rules in the bimodal risk determination engine, determining the behavior risk score data according to the user operation data; and obtaining the user's risk score data through the content risk score data and the behavior risk score data.
[0046] In this embodiment, the bimodal risk assessment engine may include behavior anomaly assessment rules and content anomaly assessment rules. The behavior anomaly assessment rules are a set of rules that determine whether abnormal behavior exists based on user operation data. Specifically, they identify potential abnormal behavior by analyzing user operation patterns, behavior frequency, operation paths, and other characteristics in the front-end system.
[0047] For example, in the behavioral anomaly judgment rule, regarding whether the terminal detects the user's operation behavior of directly submitting the file without editing, if the number of edits is 0, it can be judged as a machine operation; for example, in the behavioral anomaly judgment rule, regarding the mouse path data detected by the terminal, its frequency characteristics are analyzed through Fourier transform. If the Fourier main frequency is greater than the preset value (such as 2Hz), it can be determined that there is a high-frequency repetitive pattern in the mouse path, which is more consistent with the characteristics of machine operation, and the possibility of machine operation is high; if the Fourier main frequency is not greater than the preset value, it means that the mouse path is relatively smooth and has no obvious repetitive pattern, which is more consistent with the characteristics of human operation, and the possibility of machine operation is low.
[0048] For example, in the content anomaly judgment rules for image type files, if the terminal detects an edge density of less than 5%, it is determined that there is a high possibility that the file is generated by AI; in the content anomaly judgment rules for text type files, if the terminal detects an abnormal punctuation rate >40%, it is determined that there is a high possibility that the file is generated by a machine.
[0049] The embodiments of this specification can determine behavioral risk score data based on user operation data based on the behavioral anomaly determination rules in the dual-modal risk determination engine. The embodiments of this specification can also determine content risk score data based on content characteristics based on the content anomaly determination rules in the dual-modal risk determination engine, and then obtain the user's risk score data based on the content risk score data and the behavioral risk score data.
[0050] 230. When the risk score data meets the first preset condition, the user file information is determined to be non-AI generated content, and the behavioral fingerprint is added to the metadata corresponding to the user file information to obtain the behavioral fingerprint metadata.
[0051] In this embodiment, the first preset condition can be set as the behavior risk score data is less than the behavior risk score preset value, the content risk score data is less than the content risk score preset value, or the sum of the content risk score data and the behavior risk score data is less than the total score preset value.
[0052] In this embodiment, metadata refers to data about textual information data, providing descriptive information about the data. Metadata can include information such as the data's source, format, creation time, author, and copyright. In this embodiment, behavioral fingerprints can be added to the metadata corresponding to user file information, providing additional verification information for the file and helping to identify the file's authenticity and source.
[0053] In some embodiments, the front-end real-time identification method of AI-generated content based on bimodal verification also includes: when the risk score data meets the second preset condition, the user file information is determined to be AI-generated content, and the user file information is intercepted and processed.
[0054] In this embodiment, the second preset condition can be set as the behavior risk score data is not less than the behavior risk score preset value, the content risk score data is not less than the content risk score preset value, and the sum of the content risk score data and the behavior risk score data is not less than the total score preset value.
[0055] The embodiments of this specification adopt a dual-modal verification method of user behavior and content, comprehensively considering user operating habits and file content characteristics, thereby improving detection accuracy. It is difficult for machine-generated content to simultaneously forge human operating habits and natural content characteristics. Therefore, the embodiments of this specification can more effectively identify AI-generated content through dual-modal verification and reduce the misjudgment rate.
[0056] Moreover, the embodiments of this specification run on the front end without relying on large AI models, and achieve lightweight detection through a rule engine, which greatly reduces computing resource consumption and avoids network delays and data transmission costs caused by cloud computing.
[0057] Furthermore, the detection latency of the embodiments of this specification is reduced from 1200ms in cloud-based detection to 80ms, and the response time is less than 100ms, enabling rapid, real-time interception. This eliminates network transmission and cloud-based queuing time, allowing users to upload images or text without noticeable lag, significantly improving the user experience. This is particularly suitable for applications requiring high real-time performance, such as content publishing on social media platforms.
[0058] Furthermore, the embodiments of this specification can be run locally, with user behavior and file content processed locally without uploading to the cloud, effectively avoiding the risk of data leakage. By relying on the browser's native API for operation, it complies with privacy regulations such as GDPR, fully protecting user privacy and data security.
[0059] In addition, the embodiments of this specification comprehensively analyze user behavior and content characteristics, so that attackers need to simultaneously forge content characteristics and simulate human operation behavior, greatly increasing the cost of forgery. Uploaders of AI-generated content typically present the characteristics of "high content quality + low operation complexity", while real users often have multiple adjustments and random operation paths. The design of the present invention can effectively identify this difference, thereby enhancing the ability to identify AI-generated content and reducing the risk of being bypassed.
[0060] The embodiments of this specification use innovative dual-modal verification technology to solve the problems of single detection dimension, high computational cost, high privacy risk and poor real-time performance in the existing technology, and provide an efficient, accurate, secure and real-time solution for the recognition of AI-generated content, which has broad application prospects and important practical significance.
[0061] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0062] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a front-end real-time recognition system for AI-generated content based on dual-modal verification provided in an embodiment of this specification.
[0063] like Figure 5 As shown, the AI-generated content front-end real-time recognition system based on bimodal verification may include at least: The data acquisition module 500 is used to acquire user file information and user operation data based on the front-end environment running on the user device, and generate a behavior fingerprint based on the user operation data; Feature extraction module 510, used to perform edge density processing and word frequency jump detection processing on user file information to obtain content features of the user file information; The risk determination module 520 is configured to determine the user's risk score data based on the content characteristics and user operation data based on a preset dual-modal risk determination engine; The metadata marking module 530 is used to determine the user file information as non-AI generated content when the risk score data meets the first preset condition, and add the behavioral fingerprint to the metadata corresponding to the user file information to obtain behavioral fingerprint metadata.
[0064] In some embodiments, the data acquisition module 500 includes a fingerprint generation module, which is used to: obtain a random projection matrix, where any element in the random projection matrix is obtained by random sampling from a standard normal distribution; determine a behavior data vector through user operation data; multiply the behavior data vector with the random projection matrix to obtain a low-dimensional projection vector; determine the hash value of any dimensional element in the low-dimensional projection vector based on a hash determination engine; and generate a behavior fingerprint based on the hash value of any dimensional element in the low-dimensional projection vector.
[0065] In some embodiments, the data acquisition module 500 includes a hash determination module, which is used to: when the dimension element is less than a first preset value, the hash value corresponding to any dimension element is a first value; when the dimension element is not less than the first preset value, the hash value corresponding to any dimension element is a second value.
[0066] In some embodiments, content features include image features and text features, and the feature extraction module 510 includes a feature extraction submodule, which is used to: obtain image information and text information of user file information; perform size processing on the image in the image information to obtain an image with a preset size; calculate the edge density of the image with the preset size, and the edge density is the image feature corresponding to the user file information; obtain the abnormal punctuation rate corresponding to the text information; perform word frequency jump detection processing on the text information to obtain word frequency jump detection data; determine the context coherence score based on the word frequency jump detection data; the abnormal punctuation rate and the context coherence score are the text features corresponding to the user file information.
[0067] In some embodiments, the feature extraction submodule includes an edge density module, which is used to obtain the number of edge pixels and the total number of pixels of an image with a preset size, and the edge density is the ratio of the number of edge pixels to the total number of pixels.
[0068] In some embodiments, the feature extraction submodule includes a sliding window module, which is used to: based on a preset sliding window, slide the sliding window to process the text information, and determine the number of occurrences of each word in each sliding window; determine the word frequency change of each word in the adjacent sliding window according to the number of occurrences of each word; and obtain word frequency jump detection data through the word frequency change.
[0069] In some embodiments, the risk determination module 520 includes a scoring module, which is used to: determine content risk scoring data based on content characteristics based on content anomaly determination rules in the bimodal risk determination engine; determine behavior risk scoring data based on user operation data based on behavior anomaly determination rules in the bimodal risk determination engine; and obtain user risk scoring data through content risk scoring data and behavior risk scoring data.
[0070] In some embodiments, the AI-generated content front-end real-time identification system based on bimodal verification also includes an interception module, which is used to: when the risk score data meets the second preset condition, determine the user file information as AI-generated content and intercept the user file information.
[0071] In some embodiments, the user operation data includes the coordinates and timing of the mouse movement path, the number of file edits, and the operation time.
[0072] The embodiments of this specification use innovative dual-modal verification technology to solve the problems of single detection dimension, high computational cost, high privacy risk and poor real-time performance in the existing technology, and provide an efficient, accurate, secure and real-time solution for the recognition of AI-generated content, which has broad application prospects and important practical significance.
[0073] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the AI-generated content front-end real-time recognition system based on dual-modal verification, since it is basically similar to the embodiment of the AI-generated content front-end real-time recognition method based on dual-modal verification, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0074] The embodiment of this specification also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a computer or a processor, the computer or processor executes the above-mentioned Figures 2 to 4 If the components of the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0075] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. The technical features of this embodiment and the implementation scheme can be combined in any manner unless they conflict.
[0076] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims of this specification.
Claims
1. A front-end real-time identification method for AI-generated content based on dual-modal verification, characterized by: include: Based on the front-end environment running on the user's device, obtain user file information and user operation data, and generate a behavioral fingerprint based on the user operation data; Performing edge density processing and word frequency jump detection processing on the user file information to obtain content features of the user file information; Based on a preset dual-modal risk determination engine, determine the user's risk score data according to the content characteristics and the user operation data; When the risk score data meets the first preset condition, the user file information is determined to be non-AI generated content, and the behavioral fingerprint is added to the metadata corresponding to the user file information to obtain behavioral fingerprint metadata.
2. The method for real-time identification of AI-generated content front-end based on dual-modal verification according to claim 1, wherein the behavioral fingerprint is generated according to the user operation data, characterized in that: include: Obtain a random projection matrix, where any element in the random projection matrix is obtained by random sampling from a standard normal distribution; Determine a behavior data vector based on the user operation data; Multiplying the behavior data vector by the random projection matrix to obtain a low-dimensional projection vector; Determine the hash value of any dimensional element in the low-dimensional projection vector based on a hash determination engine; A behavioral fingerprint is generated according to a hash value of an element of any dimension in the low-dimensional projection vector.
3. The method for real-time identification of AI-generated content based on dual-modal verification according to claim 2 is characterized in that: The determining of the hash value of any dimensional element in the low-dimensional projection vector based on the hash determination engine includes: When the dimension element is less than a first preset value, the hash value corresponding to the arbitrary dimension element is a first value; when the dimension element is not less than the first preset value, the hash value corresponding to the arbitrary dimension element is a second value.
4. The method for real-time identification of AI-generated content based on dual-modal verification according to claim 1 is characterized in that: The content features include image features and text features. The edge density processing and word frequency jump detection processing are performed on the user file information to obtain the content features of the user file information, including: Acquire image information and text information of the user file information; Performing size processing on the image in the image information to obtain an image with a preset size; Calculating an edge density of the image having a preset size, where the edge density is an image feature corresponding to the user file information; Obtaining an abnormal punctuation rate corresponding to the text information; Performing word frequency jump detection processing on the text information to obtain word frequency jump detection data; A context coherence score is determined based on word frequency jump detection data; the abnormal punctuation rate and the context coherence score are text features corresponding to the user file information.
5. The method for front-end real-time identification of AI-generated content based on bimodal verification according to claim 4 is characterized in that: The calculating the edge density of the image having the preset size includes: The number of edge pixels and the total number of pixels of the image with a preset size are obtained, and the edge density is the ratio of the number of edge pixels to the total number of pixels.
6. The method for front-end real-time identification of AI-generated content based on bimodal verification according to claim 4 is characterized in that: The performing word frequency jump detection processing on the text information to obtain word frequency jump detection data includes: Based on a preset sliding window, sliding the sliding window over the text information, and determining the number of occurrences of each word in each sliding window; Determining a word frequency change of each word in adjacent sliding windows according to the number of occurrences of each word; The word frequency jump detection data is obtained through the word frequency change.
7. The method for front-end real-time identification of AI-generated content based on bimodal verification according to claim 1 is characterized in that: The preset dual-modal risk determination engine determines the user's risk score data according to the content features and the user operation data, including: Determining content risk score data based on the content features based on content anomaly determination rules in the bimodal risk determination engine; Determining behavior risk score data based on the user operation data based on the behavior anomaly determination rules in the bimodal risk determination engine; The risk score data of the user is obtained by using the content risk score data and the behavior risk score data.
8. The method for front-end real-time identification of AI-generated content based on bimodal verification according to claim 1 is characterized in that: The method further comprises: When the risk score data meets the second preset condition, the user file information is determined to be AI-generated content, and the user file information is intercepted.
9. The method for front-end real-time identification of AI-generated content based on bimodal verification according to claim 1 is characterized in that: The user operation data includes the coordinates and timing of the mouse movement path, the number of file edits, and the operation time.
10. AI-generated content front-end real-time recognition system based on dual-modal verification, characterized by: include: A data acquisition module, configured to acquire user file information and user operation data based on a front-end environment running on a user device, and generate a behavior fingerprint based on the user operation data; A feature extraction module is used to perform edge density processing and word frequency jump detection processing on the user file information to obtain content features of the user file information; A risk determination module, configured to determine the user's risk score data based on the content features and the user operation data based on a preset dual-modal risk determination engine; A metadata marking module is used to determine that the user file information is non-AI generated content when the risk score data meets the first preset condition, and to add the behavioral fingerprint to the metadata corresponding to the user file information to obtain behavioral fingerprint metadata.
Citation Information
Patent Citations
Business data-based abnormal user generation content identification method and system
CN107256257A