Copying trap verification method and device, equipment, medium and program product

Through optical character recognition and risk rule scanning technology, the problem of inefficiency in detecting Internet content copying traps has been solved, accurate verification of copied content and risk warnings have been achieved, and user information security has been guaranteed.

CN120687705APending Publication Date: 2025-09-23INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510850888.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively detect and prevent Internet content duplication traps, which threaten user information security and business security. Manual inspections are inefficient and difficult to detect minor problems.

Method used

By obtaining user authorization, we use optical character recognition technology to take screenshots and identify text on web page images, compare the copied content with the actual content character by character, and combine code analysis tools and sensitive vocabulary to perform risk rule scanning and generate risk warnings.

Benefits of technology

It realizes the accuracy verification of copied content, timely discovers and prevents copying traps, improves information security, reduces legal and reputation risks, and improves audit efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687705A_ABST
    Figure CN120687705A_ABST
Patent Text Reader

Abstract

The invention provides a copy trap verification method which can be applied to the technical field of artificial intelligence. The copy trap verification method comprises the following steps: in response to a content copy request of a user on a webpage, obtaining copy content obtained by the user; screenshot is conducted on a webpage area where the content selected by the user is located, and a webpage image is obtained; recognizing the webpage image by adopting an optical character recognition technology, and extracting actual content needing to be copied by a user; comparing the copied content with the actual content character by character to obtain a comparison result; and verifying a copy trap based on the comparison result. The invention further provides a copy trap verification device and equipment, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more specifically to a copy trap verification method, apparatus, device, medium, and program product. Background Art

[0002] With the rapid development of internet technology, online information resources have exploded, and people conduct extensive information searches and data collection online every day. Whether writing articles and reports for daily study or work, or coding in professional fields such as computer application development, the internet provides a convenient way for us to access information. However, the internet's openness and complexity also bring numerous security risks. Among them, the problem of internet content duplication is becoming increasingly prominent, posing a serious threat to user information and business security.

[0003] In everyday information usage, people often search for relevant information online and copy it for later use. However, some web pages are prone to copy traps. For example, when building websites to share content, they employ specific algorithms to create traps, so that users see content A when browsing the webpage, but when they copy and paste it locally, the content becomes content B. This seemingly minor change actually hides significant risks. For users writing articles, the copied and pasted content may contain banned words or non-compliant content. Once published, this could lead to legal disputes or damage the reputation of individuals or companies.

[0004] In the field of computer application development, the harm caused by copy traps is even more serious. When writing code, developers often search the internet for code logic and copy relevant code snippets for use. If copy traps exist in these codes, the copied and pasted code may contain numerous security vulnerabilities. Once exploited by hackers, these vulnerabilities can lead to serious consequences such as system attacks, data leaks, and business interruptions, causing huge financial losses and reputational damage to enterprises.

[0005] Currently, there's no effective method for verifying and preventing the pitfalls of internet content duplication. Existing methods primarily rely on manual visual inspection, but this approach has significant limitations. For one thing, manual inspection is inefficient, and it's difficult to meticulously examine every word and sentence when faced with large amounts of duplicated content. Furthermore, minor issues, such as subtle vulnerabilities hidden in the code and subtle banned words, are often difficult to spot manually. These undetected issues can accumulate during subsequent use, ultimately leading to major safety issues.

[0006] For example, during software development, a seemingly insignificant code vulnerability may not cause significant problems in the initial stages of a system's launch. However, as the system operates and data accumulates, this vulnerability could be exploited by hackers, leading to the collapse of the entire system. Similarly, an inappropriate wording in an article may not attract attention initially, but once widely disseminated, it could spark public doubts and criticism, negatively impacting the author or the organization involved.

[0007] Therefore, in order to prevent the security risks of copying and pasting content from the source, it is necessary to design a content duplication trap detection tool. This tool can compare the copied content and scan the risk rules, timely discover and prompt potential security issues, and provide effective security protection for users. Summary of the Invention

[0008] In view of the above problems, the present application provides a copy trap verification method, apparatus, device, medium and program product.

[0009] According to the first aspect of the present application, a copy trap verification method is provided, including: obtaining user authorization to obtain the copied content and web page images obtained by the user; after obtaining the authorization to obtain the copied content and web page images obtained by the user, in response to the user's request to copy the content of the web page, obtaining the copied content obtained by the user; taking a screenshot of the web page area where the content selected by the user is located to obtain a web page image; using optical character recognition technology to recognize the web page image and extract the actual content that the user needs to copy; comparing the copied content with the actual content character by character to obtain a comparison result; and verifying the copy trap based on the comparison result.

[0010] According to an embodiment of the present application, obtaining the copied content obtained by the user includes: monitoring the operating system clipboard and reading the clipboard content when a copy operation is detected; or directly obtaining the content to be copied to the clipboard by the browser through a browser extension interface.

[0011] According to an embodiment of the present application, a screenshot is taken of the web page area where the user-selected content is located to obtain a web page image, including: calling the operating system screenshot interface to capture the current web page window image; or locating the user-selected area through the browser node coordinates, calling the browser's built-in screenshot function, and taking an accurate screenshot of the area on the web page where the user-selected content is located.

[0012] According to an embodiment of the present application, optical character recognition technology is used to identify web page images and extract the actual content that the user needs to copy, including: preprocessing the web page image to obtain a preprocessed web page image, wherein the preprocessing includes denoising, binarization and tilt correction; performing text recognition on the preprocessed web page image, extracting text features in the web page image and performing character classification to obtain text features; and post-processing the text features to obtain the actual content that the user needs to copy, wherein the post-processing includes removing garbled characters and segmenting and dividing the text features into lines according to the layout information of the web page content.

[0013] According to an embodiment of the present application, text recognition is performed on a preprocessed web page image, text features in the web page image are extracted and character classification is performed to obtain text features, including: using a contour-based feature extraction method or a projection-based feature extraction method to extract the shape, structure and strokes of the text from the preprocessed image to obtain text features; and using a template matching method or a neural network method to match the extracted text features with a pre-trained character model to determine the specific category of each character.

[0014] According to an embodiment of the present application, based on the comparison results, the copy trap is checked, including: in response to an inconsistent comparison result, determining that a copy trap exists and preventing the spread of the copied content; and in response to a consistent comparison result, scanning the copied content for risk rules, and in response to scanning content that violates the risk rules, determining that a copy trap exists and generating a risk prompt.

[0015] According to an embodiment of the present application, the method also includes: generating a visual difference report, wherein the visual difference report includes highlighted positioning of inconsistent characters and a comparison view of tampered content; or generating a risk scanning report, wherein the risk scanning report includes code vulnerability locations and / or sensitive word trigger locations.

[0016] According to an embodiment of the present application, risk rule scanning is performed on the copied content, including: for code-type copied content, using code analysis tools or large model technology to analyze the logical structure and grammatical rules of the code, and checking whether there are code dead loops or unsafe third-party dependency calls in the code-type copied content; and for text-type copied content, using a pre-built sensitive word library and banned dictionary and a string matching algorithm to check whether the text-type copied content contains sensitive words or banned content; generating risk warnings includes: during the risk rule scanning process, recording and classifying the scanning results, and generating corresponding risk warning information according to different risk types.

[0017] The second aspect of the present application provides a copy trap verification device, including: an authorization acquisition module, used to obtain the user's authorization to obtain the copied content and web page images obtained by the user; a first acquisition module, used to obtain the copied content obtained by the user in response to the user's request to copy the content of the web page after obtaining the authorization to obtain the copied content and web page images obtained by the user; a second acquisition module, used to take a screenshot of the web page area where the content selected by the user is located to obtain the web page image; a third acquisition module, used to use optical character recognition technology to identify the web page image and extract the actual content that the user needs to copy; a comparison module, used to compare the copied content with the actual content character by character to obtain a comparison result; and a verification module, used to verify the copy trap based on the comparison result.

[0018] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0019] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0020] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0022] Figure 1 Schematically illustrates an application scenario diagram of the copy trap verification method, apparatus, device, medium, and program product according to an embodiment of the present application;

[0023] Figure 2 A flowchart of a copy trap verification method according to an embodiment of the present application is schematically shown;

[0024] Figure 3 A block diagram schematically illustrates a method for verifying a duplicate trap according to an embodiment of the present application;

[0025] Figure 4 A flowchart schematically illustrates a method of using optical character recognition technology to recognize web page images and extract the actual content that a user needs to copy according to an embodiment of the present application;

[0026] Figure 5A flowchart of risk rule scanning of duplicate content according to an embodiment of the present application is schematically shown;

[0027] Figure 6 A schematic diagram of a structure of a copy trap detection device according to an embodiment of the present application is shown; and

[0028] Figure 7 A block diagram of an electronic device suitable for implementing a copy trap checking method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0029] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0030] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0032] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0033] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0034] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided in the embodiments of the present application all provide users with corresponding operation portals for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge and skills, and have reached a certain level of professionalism.

[0035] When browsing the web, users often need to copy web content. However, some malicious web pages may set up copy traps, such as adding malicious code to the copied content or tampering with the copied content. To ensure user safety and the accuracy of the content obtained, it is necessary to accurately obtain the copied content obtained by users for subsequent copy trap verification.

[0036] An embodiment of the present application provides a copy trap verification method, including: obtaining user authorization to obtain the copied content and web page images obtained by the user; after obtaining the authorization to obtain the copied content and web page images obtained by the user, responding to the user's request to copy the content of the web page, obtaining the copied content obtained by the user; taking a screenshot of the web page area where the content selected by the user is located to obtain a web page image; using optical character recognition technology to recognize the web page image and extract the actual content that the user needs to copy; comparing the copied content with the actual content character by character to obtain a comparison result; and verifying the copy trap based on the comparison result.

[0037] Figure 1 The following schematically illustrates an application scenario of a duplication trap verification method according to an embodiment of the present application.

[0038] like Figure 1As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0039] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0040] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0041] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0042] It should be noted that the duplicate trap verification method provided in the embodiments of the present application can generally be executed by the server 105. Accordingly, the duplicate trap verification device provided in the embodiments of the present application can generally be set in the server 105. The duplicate trap verification method provided in the embodiments of the present application can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the duplicate trap verification device provided in the embodiments of the present application can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0043] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0044] The following will be based on Figure 1 The scene described by Figures 2 to 5 The copy trap verification method according to the embodiment of the present application is described in detail.

[0045] Figure 2 A flowchart of a copy trap verification method according to an embodiment of the present application is schematically shown; Figure 3 A block diagram schematically illustrates a method for verifying a duplicate trap according to an embodiment of the present application.

[0046] like Figure 2 and Figure 3 As shown, the copy trap verification method of this embodiment includes operations S210 to S260.

[0047] In operation S210, the user's authorization to obtain the copied content and web page image obtained by the user is obtained.

[0048] In operation S220, after obtaining authorization to obtain the copied content and web page image obtained by the user, in response to the user's request for copying the content of the web page, the copied content obtained by the user is obtained.

[0049] In the embodiment of the present application, the following two methods can be used to obtain the copied content obtained by the user.

[0050] The first method is to obtain the copied content by monitoring the operating system clipboard. When a user initiates a copy operation on a web page, the system temporarily stores the content to be copied in the clipboard. By monitoring changes in the clipboard, once a copy operation is detected, the content in the clipboard can be immediately read. The advantage of this method is its high versatility. It is applicable to almost all types of web pages and browsers, and does not require adaptation to a specific browser. It can obtain the user's copied content in a timely and accurate manner, regardless of whether the user triggers the copy operation through shortcut keys, menu options, or other methods. Moreover, due to direct interaction with the operating system, the obtained content is relatively complete and is not limited by the browser's internal processing logic. This facilitates subsequent accurate analysis and verification of the copied content, improves the accuracy and reliability of copy trap verification, and avoids misjudgments or omissions due to incomplete content acquisition.

[0051] The second method is to use the browser extension interface to directly obtain the content to be copied to the clipboard by the browser. The browser extension can be deeply integrated with the browser, and the content to be copied can be directly obtained through the extension interface at the moment the copy operation occurs. The advantage of this method is that it can obtain the original copied content more directly, avoiding the data conversion or loss problems that may exist in the clipboard. At the same time, browser extensions can be customized according to specific needs and optimized for the characteristics and copying scenarios of different web pages to improve the efficiency and accuracy of content acquisition. In addition, obtaining content through browser extensions can also achieve more refined control, such as selectively obtaining the content of specific web page elements, further improving the flexibility and pertinence of copy trap verification.

[0052] In operation S230 , a screenshot is taken of the web page area where the content selected by the user is located to obtain a web page image.

[0053] In an embodiment of the present application, the following two methods can be used to take a screenshot of the web page area where the content selected by the user is located to obtain a web page image.

[0054] The first method is to call the operating system's screenshot interface to capture an image of the current web page window. This method is simple to use and highly versatile, quickly capturing the current state of a web page window, including layout, style, and other information, without requiring in-depth knowledge of the page's internal structure. For simple web pages with a strong correlation between the user-selected area and the entire window, it can efficiently capture an image containing the selected content. The screenshot quality is stable, effectively restoring the web page's display, and providing the foundation for subsequent optical character recognition (OCR).

[0055] The second method is to locate the user's selected area through the browser node coordinates, and call the browser's built-in screenshot function to take an accurate screenshot. Specifically, first obtain the node of the user's selected content in the web page document object model (DOM), then obtain the node's location information, and finally call the browser's built-in screenshot function, or use the library to render the specified node area into a canvas and convert it into an image. This method is highly accurate and can ensure that the acquired image only contains the content that the user really needs to copy, avoiding interference from irrelevant information. It is especially suitable for situations where the web page content is complex and the user's selected area is small. It can reduce the amount of data for subsequent OCR processing, reduce system resource usage, and increase the response speed of screenshots and subsequent processing, thereby improving the user experience.

[0056] In operation S240 , optical character recognition technology is used to recognize the web page image and extract the actual content that the user needs to copy.

[0057] Optical character recognition technology refers to the process by which electronic devices (such as scanners, digital cameras, smartphones, etc.) determine the shape of characters by detecting the dark and light patterns in images, and then translate the character shapes into text that can be processed by computers.

[0058] Figure 4 A flowchart of using optical character recognition technology to recognize web page images and extract the actual content that the user needs to copy according to an embodiment of the present application is schematically shown.

[0059] like Figure 4 As shown, the method of this embodiment of using optical character recognition technology to recognize web page images and extracting actual content that the user needs to copy includes operations S310 to S330.

[0060] In operation S310 , the web page image is preprocessed to obtain a preprocessed web page image, wherein the preprocessing includes denoising, binarization, and tilt correction.

[0061] In the embodiments of this application, denoising aims to eliminate noise introduced during image acquisition and transmission. Common denoising methods include mean filtering and median filtering. Mean filtering replaces the center pixel value by calculating the average value of neighboring pixels, while median filtering selects the median of neighboring pixels as the new value of the center pixel. Both can effectively remove interference such as salt and pepper noise, reduce the negative impact of noise on subsequent processing, and significantly improve image quality.

[0062] Binarization is a key step in converting an image to black and white. It simplifies image information, highlights objects like text, and reduces computational complexity. The Otsu method, a global thresholding method, automatically calculates the optimal threshold based on the image's grayscale histogram, achieving rapid binarization. Local adaptive thresholding, which determines the threshold based on the characteristics of the local area surrounding a pixel, is more advantageous when processing images with uneven lighting and can better preserve target information.

[0063] Tilt correction is used to correct the tilt angle of an image, keeping objects like text horizontal or vertical. Hough transform-based methods detect straight lines in the image to determine the tilt angle and then perform rotation correction. Projection methods also accurately calculate the tilt angle and perform correction by statistically analyzing the image's horizontal and vertical projection features.

[0064] After this series of preprocessing operations, the quality of web page images is improved. Denoising removes interference, binarization makes the image information more concise and clear, and tilt correction makes the image layout regular and orderly. These changes not only speed up subsequent image processing but also significantly improve the accuracy of operations such as OCR recognition. This provides a high-quality image foundation for the entire information processing system, significantly improving system performance and efficiency.

[0065] In operation S320, perform optical character recognition (OCR) on the preprocessed web page image, extract the text features in the web page image, and classify the characters to obtain text features.

[0066] In an embodiment of the present application, when performing OCR on the preprocessed web page image, extract the text features in the web page image and conduct character classification. The obtained text features refer to the digital descriptions that can represent the essential attributes of the text and are refined from the image text information. These features cover multiple aspects. Shape features can describe geometric shape information such as the character outline, stroke direction, and thickness. For example, the letter "O" has a circular outline, and "X" is composed of two intersecting straight lines. Structural features focus on the relative positions and relationships of the internal parts of the character. For example, the Chinese character "林" consists of two "木" characters arranged side by side. Statistical features describe the character by statistically analyzing statistics such as the pixel distribution and density of the character. For example, the proportion of black pixels in the character and the average pixel gray value in the region. Transform domain features are extracted after transforming the character image to other domains (such as the frequency domain). For example, the spectral features after Fourier transform and the coefficient features after wavelet transform. These text features together constitute the digital representation of the character, laying the foundation for subsequent character classification and recognition.

[0067] Extracting these text features and performing character classification can improve the recognition accuracy. Multiple text features can more comprehensively describe the essential attributes of the character, reducing the misrecognition rate and rejection rate. It can enhance the robustness of the system. The text features can resist certain interference factors such as image noise, illumination changes, and font deformation, enabling the OCR system to maintain good performance in different environments. It helps to optimize the recognition speed. Reasonably selecting and extracting text features can reduce the computational amount, enabling the system to process a large amount of image data more efficiently. It supports multi-language recognition. The text feature extraction and classification methods have certain generality, providing convenience for cross-language information processing. It promotes subsequent text analysis. The extracted text features provide high-quality input for subsequent tasks such as text analysis, information retrieval, and machine translation, improving the performance of the entire information processing system.

[0068] In an embodiment of the present application, operation S320 includes operations S321 to S322.

[0069] In operation S321, adopt a contour-based feature extraction method or a projection-based feature extraction method to extract the shape, structure, and strokes of the text from the preprocessed image to obtain text features.

[0070] In an embodiment of the present application, in text recognition, feature extraction methods based on contours and projections can be used to extract text features from the preprocessed image. For feature extraction based on contours, first detect the text edges to obtain the contours, and then analyze features such as their shapes, perimeters, areas, convex hulls, etc. For example, calculating the convex hull of a character contour can clarify the approximate contour range, accurately capture the details of the text shape, distinguish characters with similar shapes such as "口" and "日", and can also reflect the text structure such as up and down, left and right. For feature extraction based on projections, project the image in the horizontal and vertical directions and count the pixel distribution. The horizontal projection can determine the upper and lower boundaries of characters and the line spacing, and the vertical projection can find the left and right boundaries and spacing of characters, which can quickly locate the position and distribution of characters, and judge the thickness and density of strokes. The combination of the two can comprehensively extract the shape, structure and stroke features of the text, not only can improve the text recognition accuracy, reduce noise interference, enhance the robustness of the system, but also can reduce the computational complexity, improve the processing speed, and is suitable for large-scale text recognition tasks.

[0071] In operation S322, the template matching method or the neural network method is used to match the extracted text features with the pre-trained character model to determine the specific category of each character.

[0072] In an embodiment of the present application, in the text recognition process, the template matching method or the neural network method can be used to match the extracted text features with the pre-trained character model to determine the specific category of each character.

[0073] The operation of the template matching method is relatively intuitive. First, a standard template library of characters covering different fonts, sizes and styles needs to be pre-made. During recognition, the extracted text features are compared with each template in the template library one by one, and the similarity is calculated through methods such as Euclidean distance and Hamming distance. The character category corresponding to the template with the highest similarity is the recognition result. This method is simple and direct to implement, has a small amount of calculation, fast processing speed, and low requirements for resources, and can also run on resource-constrained devices. It does not require a complex model training process. For simple and regular character recognition tasks with high real-time requirements and relatively fixed character styles, it can quickly give stable results, and is more suitable for some fixed-format bill character recognition scenarios.

[0074] Neural network principles are used to construct models such as convolutional neural networks (CNNs). During the training phase, a large number of labeled character images are fed into the network, which continuously adjusts its parameters and learns the mapping from input image features to character categories. During the recognition phase, the extracted text features are fed into the trained network, which outputs the probability of each character belonging to a different category, with the category with the highest probability being the category. This method possesses powerful learning capabilities and can automatically extract high-level image features. It is highly adaptable to complex and variable character images and can handle characters in different fonts, sizes, tilt angles, and lighting conditions. It has a high recognition accuracy rate and is particularly effective in recognizing handwritten characters and printed characters against complex backgrounds. It reduces the tediousness of manually designed features and has better generalization capabilities. It can be applied to complex scenarios such as license plate recognition and document digitization, providing a high-quality data foundation for subsequent text processing and analysis.

[0075] In operation S330 , the text features are post-processed to obtain the actual content that the user needs to copy, wherein the post-processing includes removing garbled characters and segmenting and dividing the text features into lines according to the layout information of the webpage content.

[0076] In an embodiment of the present application, after the text feature extraction is completed, it is post-processed to obtain the actual content that the user can copy. This process includes removing garbled characters, segmenting and dividing the text into lines according to the web page layout information.

[0077] To remove garbled characters, we first need to build a garbled character feature library that includes common forms of garbled characters, such as invisible characters, illegally encoded characters, and morphologically unusual symbols. We then scan the text features character by character and compare them against the feature library. If a character matches, it is identified as garbled and removed. For example, illegally encoded characters such as " " that may appear in web crawled text can be accurately identified and removed through comparison. This ensures the purity and readability of the text, prevents garbled characters from interfering with user understanding, and ensures more accurate and standardized content extraction.

[0078] The key to segmenting based on layout information is to analyze the layout tags of web pages. Tags often represent paragraphs in Hypertext Markup Language (HTML) web pages. Once the tag is recognized, the text within the tag can be divided into independent paragraphs. If there are paragraph indents or blank lines in the text, you can set rules, such as multiple consecutive spaces or a certain number of line breaks, as the basis for paragraph separation. This segmentation method can make the text structure clearer, conform to people's reading habits, and help users quickly locate and read the content of interest. It also provides a reasonable organization for subsequent text analysis and processing.

[0079] According to the layout information, pay attention to the line break mark of the web page, such as Line breaks in tags and text. When a tag or line break is used, it ends the current line and starts a new line. For example, in poetry web pages, the verses are often Tag separation, accurately identifying these tags can restore the branch structure, making it easier for users to view and use. This is especially important for texts with strict line break requirements such as poetry and code.

[0080] This series of post-processing methods can bring many beneficial effects. From the perspective of user experience, removing garbled characters and rationally segmenting and dividing the text into lines makes the text obtained by users neater and easier to read, saving time and energy for manual cleaning and improving the efficiency of information acquisition. In terms of text processing effects, it provides high-quality input for subsequent text analysis, information retrieval, machine translation and other tasks, which helps to improve the accuracy and effectiveness of these tasks. For example, in information retrieval, clearly structured text can more accurately match user query keywords. In addition, the post-processed text can adapt to diverse needs. Whether it is academic research, document editing or other application scenarios, it can better meet the needs of different users, has a wider application value, and has effectively promoted the in-depth application and development of text processing technology in various fields.

[0081] Return to reference Figure 2 and Figure 3 In operation S250, the copied content is compared with the actual content character by character to obtain a comparison result.

[0082] In the embodiments of the present application, character-by-character comparison usually adopts a sequential comparison algorithm, first determining the starting position of the two texts, and then comparing them in sequence starting from the first character. For each character position, if the corresponding character of the copied content is the same as that of the actual content, the comparison will continue to the next character; if different, the position and the difference character will be recorded. At the same time, considering that the text length may be different, when one side of the text ends and the other side has remaining characters, the remaining part is also regarded as a difference. This method is significant. It can accurately verify the accuracy of the copied content, and promptly discover errors such as missing, redundant or replaced characters, avoiding serious consequences caused by copying errors in scenarios with extremely high accuracy requirements such as legal documents and scientific research data; it can also quickly locate the error position, facilitate users to find and correct, and improve error correction efficiency; in addition, by counting the number and position of difference characters, the reliability of the copying operation can be quantitatively evaluated, providing a basis for optimizing the copying process or tools.

[0083] In operation S260 , based on the comparison result, a duplicate trap is checked.

[0084] In the embodiments of the present application, in the information copying and dissemination scenario, a copy trap detection mechanism is introduced to ensure information security and prevent the spread of malicious or inappropriate content. This mechanism first compares the copied content with the actual content character by character. If the results are inconsistent, it indicates that the copied content may have been tampered with or contains an anomaly. In this case, a copy trap is determined to exist and its dissemination is prevented. If the comparison is consistent, the copied content is further scanned for risk rules. If content that violates the risk rules is scanned, a copy trap is determined to exist and a risk warning is generated. This mechanism is effective and can intercept tampered or malicious information, ensure information security, maintain content compliance, and reduce the potential risks and losses to individuals, enterprises, and society caused by the dissemination of inappropriate content.

[0085] In the embodiments of this application, a double-check mechanism is implemented based on the comparison results. This mechanism first determines whether a copy trap exists through content comparison. If the comparison is consistent, a further risk rule scan is performed to detect potential copy traps. This double-check mechanism, on the one hand, directly determines and prevents copy traps by detecting inconsistencies, effectively preventing the spread of malicious or erroneous content and improving information security. On the other hand, if the comparison is consistent, a second check is performed through risk rule scanning, which can discover hidden content that meets specific risk rules, further enhancing the system's defense capabilities and providing users with more comprehensive security protection.

[0086] Figure 5 The flowchart of performing risk rule scanning on copied content according to an embodiment of the present application is schematically shown.

[0087] like Figure 5 As shown, the method for performing risk rule scanning on copied content in this embodiment includes operations S410 to S420.

[0088] In operation S410, the code-type copied content is analyzed using code analysis tools or large model technology to analyze the logical structure and grammatical rules of the code, and to check whether there is a code dead loop or unsafe third-party dependency call in the code-type copied content.

[0089] In the embodiments of this application, when processing code-related duplicate content, to ensure code quality and security, in-depth analysis can be performed using code analysis tools or large-scale modeling techniques. Code analysis tools, such as static code analysis tools, can perform a detailed examination of the code's logical structure and grammatical rules. They track the code execution flow and accurately determine whether there are dead loops caused by loop conditions that are always true. For example, if the condition setting in a counting loop is incorrect, the tool can promptly identify it. It also checks whether the syntax conforms to programming standards, such as bracket matching and variable definitions. It can also scan third-party dependency libraries and compare them with a list of known unsafe dependency libraries to identify unsafe calls. Large-scale modeling technology, with its powerful semantic understanding capabilities, understands the overall intent of the code, analyzes logical rationality, and detects potential dead loop risks in advance. It also combines its own knowledge to determine whether the source of the third-party dependency library is reliable and whether it has known security vulnerabilities. This method is very effective. It can not only identify code dead loops in advance, preventing the program from falling into an infinite loop during runtime, causing system resource exhaustion or crashing, but also identify unsafe third-party dependency calls, preventing security incidents such as data leakage and malicious code injection caused by dependency library vulnerabilities.

[0090] In operation S420 , for the text-type copied content, a string matching algorithm is used to check whether the text-type copied content contains sensitive words or prohibited content through a pre-built sensitive word library and prohibited content dictionary.

[0091] In the embodiments of this application, to ensure the compliance and security of copied text content, a database of sensitive terms and a prohibited dictionary covering multiple fields are constructed, and efficient string matching algorithms such as the Knuth-Morris-Pratt Algorithm and the Boyer-Moore Algorithm are used to scan and match text content word by word. This method can accurately intercept text containing sensitive or prohibited information, effectively preventing its dissemination online, on social platforms, and within the enterprise, mitigating legal risks and eliminating adverse impacts. Simultaneously, an automated detection mechanism provides risk warnings, providing a basis for subsequent audits. Its beneficial effects include maintaining a healthy online environment and reducing the spread of negative information; protecting corporate reputation and brand value; significantly improving audit efficiency and reducing labor costs, providing key technical support for cyberspace governance and corporate compliance operations.

[0092] The specific method of generating risk warnings includes: during the risk rule scanning process, recording and classifying the scanning results, and generating corresponding risk warning information according to different risk types.

[0093] In the embodiment of the present application, during the risk rule scanning process, recording and classifying the scanning results and generating corresponding risk warning information is a systematic risk management method. The specific operations are as follows: First, risk rules are set. Based on business needs, laws and regulations, and past experience, a comprehensive and detailed risk rule library is constructed. For example, in financial transaction systems, risk rules may include abnormal fluctuations in transaction amounts and frequent large-value transfers. Then, a scan is performed, using professional tools or custom scripts to conduct a comprehensive scan of target objects (such as data, systems, business processes, etc.) according to the rules. For example, a security scan of an enterprise network system is performed to check for vulnerabilities, malware, and other risks. During the scan, each risk point is recorded in detail, covering information such as the time, location, and specific content involved. For example, during a database security scan, the database table, field, and type of the vulnerability are recorded. The recorded risks are then classified according to factors such as the nature of the risk, severity, and scope of impact. They can be divided into high, medium, and low levels, or into security risks, compliance risks, business risks, etc. by type. Finally, prompt information is generated for different risk categories, including risk descriptions, possible consequences, and recommended countermeasures. For example, a high security risk prompt will say, "There is a serious vulnerability in the system that may lead to data leakage. It is recommended to immediately fix the vulnerability and strengthen security protection."

[0094] This process comprehensively and accurately identifies and locates risks within target entities, clearly defining their location and manifestations, providing a basis for subsequent action. It also promptly issues risk alerts to relevant personnel, helping them prepare in advance and prevent escalation. It also provides detailed information for management decision-making and optimizes resource allocation. Timely alerts and responses can reduce the probability and impact of risks, minimizing losses. Clear risk classification and alert information facilitate rapid risk identification and resolution, avoiding business interruptions or delays and improving operational efficiency. It also ensures business compliance and mitigates legal and reputational risks. Furthermore, this ongoing process helps refine the risk management system, accumulate experience, and enhance overall risk management capabilities.

[0095] After determining that a copy trap exists, the method further includes: generating a visual difference report, wherein the visual difference report includes highlighted locations of inconsistent characters and a comparison view of tampered content; or generating a risk scan report, wherein the risk scan report includes code vulnerability locations and / or sensitive word trigger locations.

[0096] In an embodiment of the present application, a visual difference report compares the text differences between two versions (e.g., code or documents), accurately locates and highlights inconsistent characters, and presents the tampered content in a structured view. Its core method involves using text comparison algorithms or semantic analysis techniques to visually present differences through color coding (e.g., red for additions, yellow for modifications, and strikethroughs for deletions), with arrows or annotations explaining the modification logic. This type of report is designed to quickly identify malicious changes, version conflicts, or human errors, significantly reducing manual review costs and improving efficiency, particularly in scenarios such as contract review and code merging. Its beneficial effects include: first, reducing human oversight and ensuring content consistency; second, lowering the technical barrier through visual comparison, enabling even non-professionals to quickly understand the differences; and third, providing a chain of evidence for source traceability analysis, supporting compliance audits or dispute resolution.

[0097] In the embodiments of this application, risk scanning reports focus on potential risks in code and documents, identifying code vulnerabilities (such as hard-coded passwords) and sensitive word triggers (such as personal information) through a rule engine or machine learning model. Its core methods include static code analysis, regular expression matching, and vulnerability database comparison. These methods can pinpoint the specific location of the risk (such as file path and line number) and assess the risk level. The purpose of such reports is to proactively defend against security threats, preventing data leaks or exploitation of system vulnerabilities.

[0098] In the embodiments of this application, on the one hand, by generating a visual difference report, users can intuitively and quickly identify inconsistent characters and tampered portions in copied content, helping to promptly discover and correct errors or malicious modifications, thereby improving the accuracy and security of information processing. On the other hand, the risk scan report can accurately locate code vulnerabilities and sensitive word trigger locations, providing users with detailed risk information, allowing them to take timely measures to repair and prevent them, thereby effectively reducing potential security risks. The generation of these two reports together improves the system's ability to respond to copy traps and the user's information security level.

[0099] Based on the above-mentioned duplication trap verification method, the present application also provides a duplication trap verification device. Figure 6 The device is described in detail.

[0100] Figure 6 The following schematically shows a structural block diagram of a copy trap checking device according to an embodiment of the present application.

[0101] like Figure 6 As shown, the copy trap verification device 800 of this embodiment includes an authorization acquisition module 810 , a first acquisition module 820 , a second acquisition module 830 , a third acquisition module 840 , a comparison module 850 and a verification module 860 .

[0102] The authorization acquisition module 810 is used to obtain the user's authorization to obtain the copied content and web page images obtained by the user. In one embodiment, the authorization acquisition module 810 can be used to perform the operation S210 described above, which will not be repeated here.

[0103] The first acquisition module 820 is used to obtain the copied content and web page images obtained by the user in response to the user's request for copying the content of the web page after obtaining authorization to obtain the copied content and web page images obtained by the user. In one embodiment, the first acquisition module 820 can be used to perform the operation S220 described above, which will not be repeated here.

[0104] The second acquisition module 830 is used to take a screenshot of the web page area where the user-selected content is located to obtain a web page image. In one embodiment, the second acquisition module 830 can be used to perform the operation S230 described above, which will not be repeated here.

[0105] The third acquisition module 840 is used to use optical character recognition technology to identify web page images and extract the actual content that the user needs to copy. In one embodiment, the third acquisition module 840 can be used to perform the operation S240 described above, which will not be repeated here.

[0106] The comparison module 850 is used to compare the copied content with the actual content character by character to obtain a comparison result. In one embodiment, the comparison module 850 can be used to perform the operation S250 described above, which will not be repeated here.

[0107] The verification module 860 is used to verify the duplication trap based on the comparison result. In one embodiment, the verification module 860 can be used to perform the operation S260 described above, which will not be repeated here.

[0108] According to embodiments of the present application, any multiple modules among the authorization acquisition module 810, the first acquisition module 820, the second acquisition module 830, the third acquisition module 840, the comparison module 850, and the verification module 860 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the authorization acquisition module 810, the first acquisition module 820, the second acquisition module 830, the third acquisition module 840, the comparison module 850, and the verification module 860 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the authorization acquisition module 810, the first acquisition module 820, the second acquisition module 830, the third acquisition module 840, the comparison module 850 and the verification module 860 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0109] Figure 7 A block diagram of an electronic device suitable for implementing a copy trap checking method according to an embodiment of the present application is schematically shown.

[0110] like Figure 7 As shown, an electronic device 900 according to an embodiment of the present application includes a processor 901, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 902 or programs loaded from a storage unit 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.

[0111] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in one or more memories.

[0112] According to an embodiment of the present application, electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to bus 904. Electronic device 900 may also include one or more of the following components connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 908 including a hard disk; and a communication section 909 including a network interface card such as a LAN card or modem. Communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 910 as needed, so that computer programs read from the removable media can be installed into storage section 908 as needed.

[0113] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.

[0114] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.

[0115] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to enable the computer system to implement the copy trap verification method provided in the embodiments of the present application.

[0116] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the processor 901 executes the computer program. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0117] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0118] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0119] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, Python, "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0121] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

Claims

1. A copy trap verification method, characterized in that: The method comprises: Obtaining the user's authorization to obtain the copied content and web page images obtained by the user; After obtaining authorization to obtain the copied content and web page images obtained by the user, in response to the user's request to copy the content of the web page, obtain the copied content obtained by the user; Taking a screenshot of the web page area where the user-selected content is located to obtain a web page image; Using optical character recognition technology to identify the web page image and extract the actual content that the user needs to copy; Comparing the copied content with the actual content character by character to obtain a comparison result; and Based on the comparison result, a duplicate trap is checked.

2. The method according to claim 1, characterized in that The obtaining of the copied content obtained by the user includes: By monitoring the operating system clipboard and reading the clipboard contents when a copy operation is detected; or Through the browser extension interface, directly obtain the content to be copied to the clipboard of the browser.

3. The method according to claim 1, characterized in that The method of taking a screenshot of the web page area where the content selected by the user is located to obtain a web page image includes: Call the operating system screenshot interface to capture the current web page window image; or The user's selected area is located through the browser node coordinates, and the browser's built-in screenshot function is called to take an accurate screenshot of the area where the user's selected content is located on the web page.

4. The method according to claim 1, wherein The method of using optical character recognition technology to identify the web page image and extracting the actual content that the user needs to copy includes: Preprocessing the web page image to obtain a preprocessed web page image, wherein the preprocessing includes denoising, binarization, and tilt correction; Performing text recognition on the pre-processed web page image, extracting text features from the web page image and performing character classification to obtain text features; and The text features are post-processed to obtain actual content that the user needs to copy, wherein the post-processing includes removing garbled characters and segmenting and dividing the text features into lines according to the layout information of the web page content.

5. The method according to claim 4, characterized in that The performing text recognition on the pre-processed web page image, extracting text features in the web page image and performing character classification to obtain text features includes: Extracting the shape, structure, and strokes of the text from the preprocessed image using a contour-based feature extraction method or a projection-based feature extraction method to obtain text features; and The extracted text features are matched with pre-trained character models using a template matching method or a neural network method to determine the specific category of each character.

6. The method according to claim 1, characterized in that The checking of the duplication trap based on the comparison result includes: In response to the comparison result being inconsistent, determining that a copy trap exists and preventing the spread of the copied content; and In response to the comparison result being consistent, the duplicate content is scanned for risk rules. In response to content violating the risk rules being scanned, it is determined that a duplication trap exists and a risk prompt is generated.

7. The method according to claim 6, characterized in that After determining that a duplication trap exists, the method further includes: Generate a visual difference report, wherein the visual difference report includes a highlighted location of inconsistent characters and a comparison view of the tampered content; or Generate a risk scan report, wherein the risk scan report includes code vulnerability locations and / or sensitive word trigger locations.

8. The method according to claim 6, characterized in that The risk rule scanning of the copied content includes: For code-based duplicate content, use code analysis tools or large model technology to analyze the logical structure and grammatical rules of the code to check whether there are dead loops or unsafe third-party dependency calls in the code-based duplicate content; and For text-based copied content, we use a pre-built sensitive word library and banned word dictionary, and employ a string matching algorithm to check whether the text-based copied content contains sensitive words or banned content. Generating risk warnings includes: during the risk rule scanning process, recording and classifying the scanning results, and generating corresponding risk warning information according to different risk types.

9. A copy trap verification device, characterized in that: The device comprises: An authorization acquisition module is used to obtain the user's authorization to obtain the copied content and web page images obtained by the user; A first acquisition module is configured to, after obtaining authorization to acquire the copied content and web page images acquired by the user, respond to a user's request for copying the content of the web page and acquire the copied content acquired by the user; The second acquisition module is used to take a screenshot of the web page area where the content selected by the user is located to obtain a web page image; A third acquisition module is used to use optical character recognition technology to recognize the web page image and extract the actual content that the user needs to copy; a comparison module, configured to compare the copied content with the actual content character by character to obtain a comparison result; and A verification module is used to verify the copy trap based on the comparison result.

10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.