Interaction method capable of identifying App screenshot source and providing source link

The UI style and user information characteristics of the screenshot are extracted through image recognition technology, and a jump link is generated, which solves the problem of cumbersome screenshot source recognition in the existing technology, and realizes efficient and accurate cross-platform information traceability.

CN120356227APending Publication Date: 2025-07-22刘可心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510431811.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology cannot automatically identify the source platform for screenshots. Users need to manually judge and search across platforms, which is cumbersome and inefficient.

Method used

The UI style features in the screenshot are extracted through image recognition technology, combined with user information and text features, and jump links are generated. Convolutional neural network and object detection algorithm are used to match features, and the offline mode cache and automatic jump are supported.

Benefits of technology

It realizes automatic identification of screenshot sources and generates jump links, reduces operation steps, supports offline mode to ensure a smooth experience, and improves the efficiency and accuracy of cross-platform information traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356227A_ABST
    Figure CN120356227A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of internet, in particular to a method and a system for realizing cross-platform information source tracing by combining an image recognition technology with an interaction process, and the method comprises the following steps: receiving a screenshot image input by a user; preprocessing the screenshot image, wherein the preprocessing comprises graying, noise reduction and edge enhancement; extracting UI style features in the screenshot image, wherein the UI style features comprise a platform identifier, a navigation bar layout and a color matching scheme; matching the UI style features with a platform template in a preset database, and determining a source platform; according to the interaction method capable of identifying the App screenshot source and providing the source link, the system can automatically identify the screenshot source and generate the jump link, the operation steps are greatly reduced, in addition, the system supports an offline mode, even if no network connection exists, the system can automatically identify the screenshot source and generate the jump link, and the user can directly access the original information page of the source platform. And a matching result can be cached and automatic jumping can be carried out after the network is recovered, so that smooth experience is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and particularly relates to a method and system for realizing cross-platform information source tracing through image recognition technology combined with an interaction process. It is particularly applicable to a technical solution for identifying application program interface (UI) features, user information, and content features included in a user screenshot and generating corresponding original information links. Background Art

[0002] With the popularization of Internet applications, when users browse content on different platforms (such as social media and e-commerce platforms), they often encounter the problem that it is difficult to trace the source of screenshot information. In the prior art, the image search function is mainly based on image content similarity matching. For example:

[0003] Taobao: Search for similar products through image content, but it cannot identify the UI features in the screenshot or jump to the original product page.

[0004] Xiaohongshu: Match similar posts through image content. However, if the post content is a screenshot from another platform (such as TikTok), the user needs to manually identify the UI style and search across platforms, and the process is cumbersome.

[0005] Defects of the prior art:

[0006] Functional limitations: Existing image search only focuses on content similarity and lacks the ability to identify the source of the screenshot (such as platform UI and user information).

[0007] High operation threshold: The user needs to judge the source of the screenshot by himself, manually input keywords for search, and also needs to handle language barriers or duplicate results.

[0008] Low efficiency: Cross-platform information tracing relies on manual operations, which takes a long time and has a low success rate.

[0009] Therefore, there is an urgent need for a solution that can automatically identify the source of the screenshot and directly jump to the original information. Summary of the Invention

[0010] The purpose of the present invention is to solve the disadvantages existing in the prior art, and to propose an interaction method that can identify the source of an App screenshot and provide a source link.

[0011] To achieve the above purpose, the present invention adopts the following technical solution: An interaction method that can identify the source of an App screenshot and provide a source link, including the following steps:

[0012] Receive a screenshot image input by a user;

[0013] Preprocess the screenshot image, including grayscale conversion, noise reduction, and edge enhancement;

[0014] Extract the UI style features in the screenshot image, including platform identification, navigation bar layout, and color scheme;

[0015] Match the UI style features with the platform templates in the preset database to determine the source platform;

[0016] Generate a jump link for the user to directly access the original information page of the source platform.

[0017] Preferably, the extraction of the UI style features uses a Convolutional Neural Network (CNN) model to output a 128-dimensional feature vector, and the matching degree is calculated through cosine similarity.

[0018] Preferably, it also includes extracting user information features in the screenshot image, including:

[0019] Locate the user avatar area through the target detection algorithm;

[0020] Perform hash encoding on the avatar area to generate a 64-bit binary hash value;

[0021] Compare the hash value with the user information in the database to determine the user identity.

[0022] Preferably, it also includes extracting text information in the screenshot image, including:

[0023] Identify the user name and content text through OCR technology;

[0024] Extract keywords as search parameters and generate a jump link.

[0025] Preferably, the generation of the jump link includes calling the API interface of the target platform and passing in parameters such as the user name, content ID, or timestamp.

[0026] Preferably, if the matching fails, provide the function of manually selecting a platform and guide the user to enter the search interface of the platform.

[0027] A method for tracing the information source with multi-modal feature fusion includes the following steps:

[0028] Extract UI style features, user avatar hash features, and text keyword features from the screenshot image;

[0029] Assign weights to the features, where the UI style features account for 50%, the user avatar hash features account for 30%, and the text keyword features account for 20%;

[0030] Calculate the comprehensive score. If the score ≥ 0.7, it is determined that the matching is successful and a jump link is generated.

[0031] Preferably, it also includes dynamically updating the preset database, including:

[0032] Periodically capture the latest UI screenshots of each platform;

[0033] Modify the template data according to user feedback to ensure the matching accuracy rate.

[0034] Preferably, a method for tracing the information source in an offline mode includes the following steps:

[0035] Deploy a lightweight recognition model locally to extract and match features of the screenshot image;

[0036] Pre-load the UI templates of high-frequency platforms into the local cache;

[0037] If the network is unavailable, generate a temporary link that will automatically jump when the network is restored.

[0038] An information source tracing system includes:

[0039] An image acquisition module for receiving the screenshot image input by the user;

[0040] A preprocessing module for performing standardization processing on the image;

[0041] A feature extraction module for extracting UI styles, user information, and text features;

[0042] A matching and jumping module for generating a jump link and calling the target platform API.

[0043] The present invention has the following beneficial effects:

[0044] An interactive method designed by the present invention can identify the source of an App screenshot and provide a source link. This system can automatically identify the screenshot source and generate a jump link, greatly reducing the operation steps. For example, when a user sees a TikTok screenshot on a social media, they only need to upload the picture, and the system can quickly match it to the original video and generate a direct link, saving the cumbersome process of copying keywords and manual searching. In addition, the system supports the offline mode. Even without a network connection, it can cache the matching results and automatically jump when the network is restored, ensuring a smooth experience. Description of the Drawings

[0045] Figure 1 It is a schematic structural diagram of the system in the present invention;

[0046] Figure 2 It is a schematic diagram of an APP screenshot in the present invention;

[0047] Figure 3 It is a schematic diagram of taking pictures and recognizing pictures and recognizing pictures by the present invention;

[0048] Figure 4It is a schematic flowchart of the recognition result page in the present invention;

[0049] Figure 5 It is a schematic diagram of successful recognition in the present invention;

[0050] Figure 6 It is a schematic diagram of the jump interface in the present invention;

[0051] Figure 7 It is a schematic diagram of successful interface jump in the present invention. Detailed implementation manners

[0052] For ease of understanding the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. The preferred embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present invention can be understood more thoroughly and comprehensively.

[0053] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be a middle element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a middle element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration and do not represent the only embodiments.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0055] Embodiment 1:

[0056] An interactive method capable of identifying the source of an App screenshot and providing a source link. The core of this embodiment is to quickly match to the source platform and generate a jump link by identifying the UI style features in the user's screenshot.

[0057] The system architecture is divided into the following modules:

[0058] Image acquisition module: It supports the user to select a screenshot from the local album or take a real-time photo through the camera.

[0059] Preprocessing module: It performs standardization processing on the image, including size adjustment, noise elimination and feature enhancement.

[0060] Feature extraction module: A deep learning model is used to extract UI style, user information, and content features.

[0061] Database module: Stores UI templates for each platform, user information hash values, and content keyword indexes.

[0062] Matching and jumping module: Performs feature matching and calls the target platform API to generate jump links.

[0063] The image acquisition and preprocessing include image processing methods, image preprocessing processes, UI style feature extraction and matching, user information extraction, jump link generation and verification, exception handling and optimization, where:

[0064] Image input methods include:

[0065] Album selection: After the user enters the system, a list of local screenshots sorted in reverse chronological order is displayed by default (as Figure 2 shown), and the recognition process is entered after clicking the target image.

[0066] Real-time shooting: If the user selects to take a picture with the camera, the system calls the device camera permission and supports autofocus and image cropping functions to ensure that the key areas of the screenshot (such as the UI navigation bar and user name) are clearly visible.

[0067] The image preprocessing process includes:

[0068] Grayscale processing: Converts the RGB image to a grayscale image to reduce computational complexity. The formula is:

[0069] Gray = 0.299R + 0.587G + 0.114B

[0070] Noise reduction processing: Uses Gaussian filtering (kernel size 5×5, standard deviation σ = 1.5) to eliminate image noise.

[0071] Edge enhancement: Uses the Canny edge detection algorithm to extract the contour features of UI elements, with parameter settings of a low threshold of 50 and a high threshold of 150.

[0072] Size normalization: Scales the image to a standard size (such as 512×512 pixels) to ensure compatibility with subsequent model inputs.

[0073] UI style feature extraction and matching include:

[0074] UI template library construction: By regularly capturing typical interface screenshots of mainstream platforms (such as TikTok, Xiaohongshu, Weibo), covering different versions of UI styles (such as navigation bar layout, icon design, color scheme); then annotating the following features for each screenshot:

[0075] Platform identification: such as the "musical note" icon of TikTok and the "red book" Logo of Xiaohongshu;

[0076] Layout structure: navigation bar position (top / bottom), icon arrangement method (horizontal / vertical);

[0077] Color coding: extract the RGB value of the main color (such as the black background #000000 of TikTok);

[0078] Template storage: Store the labeled features in the database, and each record contains the platform name, feature vector, and update timestamp.

[0079] Feature matching algorithm: Use ResNet-50 as the backbone network, input the preprocessed grayscale image, output a 128-dimensional feature vector, and use cosine similarity to compare the screenshot feature with the feature vector in the template library. The formula is:

[0080]

[0081] Among them, A is the screenshot feature vector and B is the template feature vector.

[0082] Matching threshold setting: When the similarity ≥ 0.85, it is determined as a successful match, otherwise it enters the manual assistance process.

[0083] User information extraction: Locate the avatar area in the screenshot through the YOLOv5 model, set the confidence threshold to 0.7, and perform the following processing on the avatar area:

[0084] Scale to 64×64 pixels;

[0085] Convert to grayscale image;

[0086] Calculate the difference hash (dHash) to generate a 64-bit binary hash value.

[0087] Calculate the Hamming distance between the dHash value of the screenshot avatar and the hash value in the database. When the distance ≤ 5, it is determined as the same user. Use the Tesseract OCR engine to recognize the user name and post content text in the screenshot, support mixed Chinese and English recognition, and extract high-frequency words (such as "@username", "#topic tag") as search keywords.

[0088] Jump link generation and verification, design API request parameters for different platforms:

[0089] TikTok: Through the / api / user / search interface, pass in the username or video ID;

[0090] Xiaohongshu: Call the / xhs / content / get interface and pass in the post title keyword.

[0091] Parameter Encoding: URL-encode special characters (such as spaces, @) to avoid request errors. After successful matching, a prompt box will pop up on the interface (as Figure 4 shown), displaying the source platform name and a jump button.

[0092] If the user has not installed the target platform application, they will be guided to the application store download page.

[0093] Exception Handling and Optimization

[0094] Failure Prompt: If the matching similarity < 0.85, the interface will display "Recognition failed. It is recommended to reshoot or manually select the platform";

[0095] Manual Selection: Provide a platform list (such as TikTok, Xiaohongshu, Weibo). After the user clicks, they will enter the search interface of that platform.

[0096] Example 2: Multimodal Feature Fusion Recognition

[0097] Multifaceted Feature Fusion Architecture

[0098] In this example, by combining multimodal features of UI styles, user avatars, and text content, the recognition accuracy in complex scenarios is improved.

[0099] Feature Weight Allocation

[0100] UI Style Feature: Weight ratio 50% (decisive factor);

[0101] User Avatar Hash: Weight ratio 30%;

[0102] Text Keywords: Weight ratio 20%.

[0103] Fusion Formula

[0104] Comprehensive Score = 0.5 × S UI + 0.3 × S 头像 + 0.2 × S 文本

[0105] where S UI is the UI style similarity, S 头像 is the avatar hash matching degree, and S 文本 is the keyword matching degree.

[0106] Dynamic Database Update: Regularly run the Headless Chrome browser to simulate user access to each platform page and capture the latest UI styles;

[0107] Version Control: Maintain a version history record for each platform to avoid interference from old templates in new version recognition.

[0108] User feedback integration

[0109] Add a "Correction Source" button on the jump result page, where users can select the correct platform or enter keywords; filter low-quality feedback (such as incorrect operation) and only retain correction data with a confidence level ≥ 90%.

[0110] Cross-language support

[0111] Expand the language pack of the Tesseract engine to support recognition of non-Latin characters such as Japanese and Korean; use word segmentation technology for mixed-language text to extract cross-language keywords (such as "@user_日本語"); for foreign user names, call the GoogleTransliterate API to convert them into pinyin form to facilitate cross-platform search.

[0112] Embodiment 3:

[0113] The system also includes an offline mode, which includes model compression, caching and synchronization strategies.

[0114] The model compression: converts the 32-bit floating-point weights of ResNet-50 into 8-bit integers, reducing the model volume by 75%; removes convolution kernels with low contribution and retains the core feature extraction capability; uses the TensorFlow Lite framework to achieve real-time feature extraction on the mobile terminal with a delay of ≤200ms.

[0115] The cache and synchronization strategy described above is as follows: based on user history records, UI templates of the top 10 platforms are preloaded into local storage; the server is connected once every 24 hours to download incremental update packages; if the user is in an offline environment, the system generates a "temporary link" and automatically jumps after the network is restored.

[0116] When the user uploads a screenshot or real-time shooting interface, the system first performs pre-processing such as grayscale conversion, noise reduction, and edge enhancement on the image to standardize the input; then, the ResNet-50 model is used to extract UI style features (such as navigation bar layout, platform icons, and color coding), and a cosine similarity comparison is performed with the pre-built dynamically updated template library (threshold ≥ 0.85). At the same time, the avatar hash value matching located by YOLOv5 (Hamming distance ≤ 5) and the text keywords extracted by OCR are combined, and multimodal features are fused with a weight of 50%-30%-20% to generate a comprehensive matching score; after a successful match, the target platform API is called to generate a jump link (such as TikTok's video ID interface or Xiaohongshu's post search interface), and cache template matching in offline mode and delayed jump after network recovery are supported. The entire process covers image processing, feature fusion, dynamic database update, and abnormal feedback optimization, realizing closed-loop processing from screenshot recognition to precise jump.

[0117] This interactive system provides users with an efficient and accurate screenshot source identification and jump solution through intelligent image recognition and multi-modal feature matching technologies. Its core advantages are reflected in the following aspects:

[0118] Traditional screenshot sharing methods usually require users to manually search for the source, while this system can automatically identify the screenshot source and generate a jump link, significantly reducing the number of operation steps. For example, when a user sees a TikTok screenshot on a social media platform, they only need to upload the picture, and the system can quickly match it to the original video and generate a direct link, saving the cumbersome process of copying keywords and manual searching. In addition, the system supports an offline mode, which can cache the matching results even without a network connection and automatically jump after the network is restored, ensuring a smooth experience.

[0119] The system adopts a multi-modal feature fusion strategy, combining UI styles (50% weight), user avatar hashes (30% weight), and text keywords (20% weight) for comprehensive matching, avoiding misjudgments caused by single features. For example, even if screenshots from different platforms contain similar text content (such as "Popular Recommendations"), the system can still accurately distinguish the source through visual features such as the UI navigation bar layout and platform logos. At the same time, the dynamically updated UI template library can adapt to interface changes on various platforms, ensuring long-term recognition accuracy.

[0120] The system pre-installs UI templates and API interfaces for mainstream platforms such as TikTok, Xiaohongshu, and Weibo, which can cover the needs of the vast majority of users. For niche platforms that are not included, users can supplement data through manual selection or feedback mechanisms to gradually expand the recognition scope. In addition, the system supports multi-language OCR recognition (such as Chinese, English, Japanese, and Korean), and can convert foreign usernames into pinyin to improve the accuracy of cross-language searches.

[0121] The system adopts de-identification processing technology. The screenshots uploaded by users are only used for feature extraction (such as generating avatar hash values and OCR text keywords), and the original images will not be stored for a long time. During the matching process, user information is transmitted in an encrypted form to avoid the leakage of sensitive data. At the same time, the system supports local processing, and key calculations (such as ResNet-50 feature extraction) can be completed on the device side, reducing dependence on the cloud and further protecting privacy.

[0122] Through model compression technologies (such as 8-bit integer quantization and removing redundant convolutional kernels), the system achieves millisecond-level response (delay ≤ 200ms) on mobile devices. Gaussian filtering and Canny edge detection in the preprocessing stage further accelerate feature extraction, enabling smooth operation even on low-end devices.

[0123] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. An interactive method capable of identifying the source of an App screenshot and providing a source link, characterized in that It includes the following steps: Receive the screenshot image input by the user; Preprocess the screenshot image, including grayscale conversion, noise reduction, and edge enhancement; Extract the UI style features in the screenshot image, including platform identification, navigation bar layout, and color scheme; Match the UI style features with the platform templates in the preset database to determine the source platform; Generate a jump link for the user to directly access the original information page of the source platform.

2. The interactive method for identifying the source of an App screenshot and providing a source link according to claim 1, wherein, The extraction of the UI style features uses a convolutional neural network (CNN) model to output a 128-dimensional feature vector, and the matching degree is calculated through cosine similarity.

3. The interactive method for identifying the source of an App screenshot and providing a source link according to claim 1, characterized in that It also includes extracting the user information features in the screenshot image, including: Locate the user avatar area through the target detection algorithm; Perform hash encoding on the avatar area to generate a 64-bit binary hash value; Compare the hash value with the user information in the database to determine the user identity.

4. An interactive method capable of identifying the source of an App screenshot and providing a source link according to claim 1, characterized in that, It also includes extracting the text information in the screenshot image, including: Identify the user name and content text through OCR technology; Extract keywords as search parameters and generate a jump link.

5. The interactive method for identifying the source of an App screenshot and providing a source link according to claim 1, wherein The generation of the jump link includes calling the API interface of the target platform and passing in parameters such as user name, content ID, or timestamp.

6. The interactive method for identifying the source of an App screenshot and providing a source link according to claim 1, characterized in that If the matching fails, provide the function of manually selecting a platform and guide the user to enter the search interface of the platform.

7. A method for tracing the information source of multimodal feature fusion, characterized in that, It includes the following steps: Extract the UI style features, user avatar hash features, and text keyword features from the screenshot image; Assign weights to the features, where the UI style features account for 50%, the user avatar hash features account for 30%, and the text keyword features account for 20%; Calculate the comprehensive score. If the score ≥ 0.7, it is determined that the matching is successful and a jump link is generated.

8. The method according to claim 7, wherein It also includes dynamically updating the preset database, including: Regularly capture the latest UI screenshots of each platform; Correct the template data according to user feedback to ensure the matching accuracy.

9. An information source tracing method in an offline mode, characterized in that, It includes the following steps: Deploy a lightweight recognition model locally to extract and match features from the screenshot image; Preload the UI templates of high-frequency platforms into the local cache; If the network is unavailable, generate a temporary link and automatically jump after the network is restored.

10. An information source tracing system, characterized in that, It includes: An image acquisition module for receiving the screenshot image input by the user; A preprocessing module for performing standardized processing on the image; A feature extraction module for extracting UI styles, user information, and text features; A matching and jump module for generating a jump link and calling the target platform API.