A cross-end copy method and related apparatus

By using camera capture and multimodal recognition technology, cross-device data copying under network-free conditions is realized, solving the convenience and compatibility issues of cross-device copying in existing technologies. It supports cross-device copying of various information types, improving operational convenience and compatibility.

CN121597446BActive Publication Date: 2026-06-02BEIJING SOHU INTERNET INFORMATION SERVICE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SOHU INTERNET INFORMATION SERVICE
Filing Date
2026-01-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing cross-device copying technologies cannot achieve convenient and cross-ecosystem compatible physical information collection and cross-device copying in offline environments, and also pose issues of data privacy leakage and ecosystem closure.

Method used

By capturing and copying images using a camera for pagination, and utilizing multimodal recognition and intelligent bounding box technology, it achieves structured processing and copying of multimodal data such as text, images, and tables, and supports seamless cross-platform copying.

Benefits of technology

It enables real-time, convenient, and compatible copying of cross-device data even without a network connection, breaking through the environmental and ecological limitations of traditional technologies, supporting cross-device copying of various information types, and improving operational convenience and compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597446B_ABST
    Figure CN121597446B_ABST
Patent Text Reader

Abstract

The application provides a cross-terminal copying method and related device, and relates to the technical field of software, which comprises the following steps: capturing a copy picture through a camera and performing pagination processing on the copy picture to obtain at least one single page; performing multi-modal identification on each single page to obtain structured data of each single page and outputting the structured data, wherein the structured data comprises multi-modal data in a multi-modal area of the single page; determining target multi-modal data framed in at least one single page, and copying the target multi-modal data to a target editing position. The application can realize seamless copying function with a target application, complete direct copying from a physical space to a digital space without a network, keep operation convenient, and ensure real-time performance and compatibility of cross-terminal data transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software technology, and in particular to a cross-platform copying method and related apparatus. Background Technology

[0002] With the accelerating evolution of mobile office and multi-device collaboration trends, users' demand for cross-terminal information interaction has exploded, and cross-device copying of multimodal data can significantly improve the efficiency of multi-device workflows.

[0003] Therefore, how to achieve cross-device copying that is not network-required, easy to operate, compatible with various ecosystem devices, and supports the collection of physical information has become an urgent problem to be solved at this stage. Summary of the Invention

[0004] In view of the above problems, this application provides a cross-device copying method and related apparatus to achieve cross-device copying that is network-free, easy to operate, compatible with various devices across different ecosystems, and supports the collection of physical information. The specific solution is as follows:

[0005] A first aspect of this application provides a cross-end copying method, the cross-end copying method comprising:

[0006] The copied image is captured by a camera, and the copied image is paginated to obtain at least one single page;

[0007] Multimodal recognition is performed on each single page to obtain the structured data of each single page and output it. The structured data includes the multimodal data within the multimodal region of the single page.

[0008] Identify the target multimodal data that is selected within the frame on at least one single page, and copy the target multimodal data to the target editing location.

[0009] In one possible implementation, the paging process of the copied screen to obtain at least one single page includes:

[0010] If the copied image is a single image, the copied image is denoised and cropped to obtain the target image;

[0011] The target screen is subjected to page detection, and at least one detected page area is cropped to obtain at least one single page;

[0012] The at least one single page is deduplicated based on perceptual hashing, feature matching, and OCR recognition.

[0013] In one possible implementation, the paging process of the copied screen to obtain at least one single page includes:

[0014] When the copied image is a video, multiple keyframes are extracted from the copied image through scene detection and layout detection.

[0015] The multiple keyframes are deduplicated based on the inter-frame difference method;

[0016] Multiple target keyframes are determined by clustering the deduplication results of the multiple keyframes, and the multiple target keyframes are stitched together. The multiple target keyframes belong to different categories, and each target keyframe is the keyframe with the highest clarity in its category.

[0017] Page detection is performed on the splicing result of the multiple target keyframes, and at least one detected page region is cropped to obtain at least one single page.

[0018] In one possible implementation, determining the target multimodal data selected within the at least one single page includes:

[0019] In response to a selection operation of a target multimodal region under a target single page, the target multimodal region is highlighted, and the multimodal data within the target multimodal region is used as the target multimodal data;

[0020] When the target multimodal region is a table region or a text region, in response to the selection operation on the target multimodal data, the selected data is filtered from the target multimodal data.

[0021] In one possible implementation, copying the target multimodal data to the target editing location includes:

[0022] The target multimodal data is stored in layers to obtain clipboard data;

[0023] A function menu is generated based on the clipboard data and the format types supported by the target editing location;

[0024] In response to the selection operation of the function menu, the target function selected in the function menu is determined, and the target data in the clipboard data that matches the target function is copied to the target editing location.

[0025] In one possible implementation, generating a function menu based on the clipboard data and the format types supported by the target editing location includes:

[0026] The lifespan of the clipboard data is obtained, and within the lifespan, the function menu is generated based on the clipboard data and the format type supported by the target editing position.

[0027] A second aspect of this application provides a cross-end copying device, the cross-end copying device comprising:

[0028] The pagination processing module is used to capture a copied image through a camera and perform pagination processing on the copied image to obtain at least one single page;

[0029] A multimodal recognition module is used to perform multimodal recognition on each single page to obtain and output the structured data of each single page. The structured data includes multimodal data within the multimodal region of the single page.

[0030] The data copy module is used to determine the target multimodal data selected in the at least one single page and copy the target multimodal data to the target editing position.

[0031] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the cross-end copying method of the first aspect or any implementation thereof.

[0032] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0033] The memory is used to store computer programs;

[0034] The processor is used to execute the computer program to enable the electronic device to implement the cross-end copying method of the first aspect or any implementation thereof.

[0035] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform a cross-end copying method as described in the first aspect or any implementation thereof.

[0036] By employing the above technical solutions, this application provides a cross-platform copying method and related apparatus, comprising: capturing a copying image using a camera and performing pagination processing on the copying image to obtain at least one single page; performing multimodal recognition on each single page to obtain structured data for each single page and outputting the structured data, wherein the structured data includes multimodal data within the multimodal region of the single page; determining the target multimodal data selected within the frame of at least one single page, and copying the target multimodal data to the target editing location. This application utilizes a device camera to achieve non-contact data acquisition, supports input of visual data such as device screens or physical documents / posters, and then precisely locates the target data to be copied through pagination, multimodal recognition, and intelligent frame selection, ultimately achieving copying at the target editing location. This application can achieve seamless copying functionality with the target application, completing direct copying from physical space to digital space without a network, ensuring real-time performance and compatibility of cross-platform data transmission while maintaining operational convenience. Attached Figure Description

[0037] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0038] Figure 1 A flowchart illustrating a cross-end copying method provided in an embodiment of this application;

[0039] Figure 2 This is a partial flowchart illustrating a cross-end copying method provided in an embodiment of this application;

[0040] Figure 3 This is another schematic flowchart of a cross-end copying method provided in an embodiment of this application;

[0041] Figure 4 This is another schematic flowchart of a cross-end copying method provided in an embodiment of this application;

[0042] Figure 5 This is another schematic flowchart of a cross-end copying method provided in an embodiment of this application;

[0043] Figure 6 This is a schematic diagram of the structure of a cross-end copying device provided in an embodiment of this application;

[0044] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0046] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0047] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0048] Seamless multimodal data copying has become a core requirement for improving the efficiency of multi-device workflows. This requirement is not only reflected in transmission speed, but also emphasizes the convenience of operation and cross-platform compatibility. Traditional information copying methods can no longer meet the requirements of modern users for compatibility, convenience and immediacy.

[0049] Existing cross-platform replication technology systems can be divided into three main categories: "cloud synchronization technology", "near-field transmission technology" and "in-application tool solutions". Each type of solution has significant differences in technical principles and application scenarios.

[0050] 1) Cloud synchronization technology. This technology enables cross-device sharing by uploading clipboard content to a cloud server. Its core reliance is on user account binding and a stable network environment. This type of solution has an average latency of 5 seconds, is highly dependent on the network and platform, has a high failure rate in weak network environments, and poses a risk of data privacy leaks.

[0051] 2) Near Field Transmission (NFC) technology. This technology enables direct communication between devices based on Bluetooth or Wi-Fi Direct protocols. Its effective transmission distance is typically limited to within 10 meters. In environments with many devices, Bluetooth transmission has a low success rate and requires users to manually trigger the connection, which is a cumbersome process. Furthermore, the copying process still requires multiple devices to have the same application installed simultaneously.

[0052] 3) In-app tool solutions require specific software to be pre-installed on both devices to achieve clipboard synchronization within the application via a proprietary protocol. However, these tools suffer from ecosystem closure issues, requiring the same application to be pre-installed on multiple devices and failing to cover system-level application scenarios.

[0053] To address the aforementioned issues, this application provides a cross-device copying method that enables information transfer between multiple devices and between devices and physical objects, achieving "seamless copying." This method is not only fast but also convenient to operate and compatible with cross-platform systems. The cross-device copying method of this application will be described in detail below with reference to the accompanying drawings.

[0054] See Figure 1 , Figure 1 This is a flowchart illustrating a cross-end copying method provided in an embodiment of this application. Figure 1 As shown in the figure, the cross-end copying method provided in this application embodiment may include steps S101 to S103, which are described in detail below.

[0055] S101, capture the copied image using a camera, and perform pagination processing on the copied image to obtain at least one single page.

[0056] In this embodiment, by activating the camera, continuous shooting or video recording can be performed on mobile phone screen content, multi-page documents, subway billboards, etc., to obtain a copy image, which can be an image or video. To adapt to cross-device copying, the copy image acquisition can support multiple scenarios (such as close-range document scanning or long-range screen capture), and can also support automatic adjustment of shooting focus or angle.

[0057] When shooting in dynamic scenes, contrast-based autofocus (AF) algorithms can be used to achieve rapid focusing by analyzing changes in high-frequency components of the image in real time. This sharpens the edges of key information such as text paragraphs or QR codes, solving the problem of clarity in dynamic imaging. Furthermore, to address image quality degradation caused by complex lighting conditions such as backlighting and shadows, a CLAHE (Contrast-Limited Adaptive Histogram Equalization) enhancement algorithm can be embedded. This algorithm improves local contrast through block processing while suppressing noise amplification, ensuring that details in dark areas are discernible.

[0058] In practical applications, image resolution adaptive adjustment technology can be used to construct a dynamic adjustment range. When the content in the copied image is dense text or a small QR code, it can automatically switch to 4K mode to retain details. For ordinary interface elements or large text, 1080P or 720P resolution is selected to balance the processing speed.

[0059] Furthermore, after obtaining the copied image, a page segmentation model (such as LayoutParser / DocBank / YOLOv8-Layout) can be used to detect the page regions in the copied image and crop each page region into a single page.

[0060] In one possible implementation, when copying a single image, denoising and then pagination can be performed, with deduplication based on image content per page. This improves the quality of single-page copying for multimodal recognition of image-based images. See also Figure 2 , Figure 2 This is a partial flowchart illustrating a cross-end copying method provided in an embodiment of this application. Figure 2 As shown in the embodiment of this application, a cross-platform copying method is provided, wherein step S101, "performing pagination on the copied screen to obtain at least one single page", may include steps S201 to S203, which are described in detail below.

[0061] S201, when the copied image is a single image, the copied image is denoised and cropped to obtain the target image.

[0062] In this embodiment of the application, when the copied image is a single image, image preprocessing can be performed by denoising and cropping. Specifically, non-local mean denoising or bilateral filtering can be used for denoising and sharpening. Canny contour detection and four-point perspective transformation are then used to detect the edges of the document page. Automatic cropping is then performed according to the edges to remove irrelevant backgrounds. Finally, Hough line detection and affine transformation are used to complete tilt correction.

[0063] S202, perform page detection on the target screen, and crop at least one detected page area to obtain at least one single page.

[0064] In this embodiment of the application, a page segmentation model (such as LayoutParser / DocBank / YOLOv8-Layout) can be used to detect page regions in the target screen, and each page region can be cropped into a single page to obtain at least one single page.

[0065] S203, deduplicating at least one single page based on perceptual hashing, feature matching, and OCR recognition.

[0066] In this embodiment, perceptual hashing is first used to calculate the pHash / dHash of each page. When the similarity between the pHash / dHash of two pages is greater than the corresponding threshold (e.g., 95%), they are determined to be duplicates, and one of the pages is deleted. Further, feature matching (e.g., SIFT to detect DoG extrema / ORB to detect corners) is used to remove duplicates from low-quality or rotated pages. Even further, OCR is used to identify the text content of each page. When the similarity between the text content of two pages is greater than the corresponding threshold, they are determined to be duplicates, and one of the pages is deleted.

[0067] In one possible implementation, when copying video footage, keyframes can be removed and clustered to determine the highest-resolution keyframe in each cluster. This clustered keyframe is then further stitched together and processed in pages, thereby improving the single-page quality of video copies in multimodal recognition. See also Figure 3 , Figure 3 This is another part of the flowchart illustrating a cross-end copying method provided in an embodiment of this application. For example... Figure 3 As shown in the embodiment of this application, a cross-platform copying method is provided, wherein step S101, "performing pagination on the copied screen to obtain at least one single page", may include steps S301 to S304, which are described in detail below.

[0068] S301 extracts multiple keyframes from the copied image when the copied image is a video image through scene detection and layout detection.

[0069] In this embodiment, when the copied image is video, scene detection (such as PySceneDetect) can be used to identify camera transitions, filtering out keyframes with a high proportion of faces or intense movement (such as non-PPT pages, document pages, and heavily blurred pages). A lightweight CNN or a layout detection model can then be trained to identify layout regions (such as slide areas, document page areas, poster areas, etc.) in the remaining keyframes and filter out keyframes containing the specified layout regions. Additionally, color / edge stability can be used to identify and filter keyframes that are typically static, high-contrast, and rectangular in shape, such as document pages and PPT pages.

[0070] S302, based on the inter-frame difference method, performs deduplication on multiple keyframes.

[0071] In this embodiment, for situations where the same PPT page or document page is viewed for an extended period, multiple keyframes obtained in the above steps can be deduplicated using the inter-frame difference method. If the change in consecutive frames is less than the corresponding threshold, the keyframe is skipped. Of course, in practical applications, perceptual hashing can be further used to calculate the pHash of each keyframe, thereby identifying two keyframes with a similarity greater than the corresponding threshold as duplicates and deleting one of them.

[0072] S303 determines multiple target keyframes by clustering the deduplication results of multiple keyframes, and stitches the multiple target keyframes together. The multiple target keyframes belong to different categories, and each target keyframe is the keyframe with the highest clarity in its category.

[0073] In this embodiment, K-means clustering is performed on the deduplication results of multiple keyframes, and Laplacian variance is used to determine the clearest frame in each keyframe class for retention. This yields multiple target keyframes, each belonging to a different class, and each keyframe being the clearest keyframe within its class. For example, if 30 keyframes appear in the same PPT within 10 seconds, only the clearest 10th keyframe is retained.

[0074] Furthermore, in scenarios such as shooting large and small advertisement pages, there are multiple consecutive frames that are different parts of the same page. In this case, the first keyframe among multiple target keyframes can be used as the starting point to determine the next keyframe with the largest overlap, and so on to construct a chain sequence of keyframes. The similarity matrix between any two keyframes can be calculated, the vertex coordinates of the image in the global coordinate system can be marked, and the clearest keyframe can be selected first for stitching.

[0075] S304, perform page detection on the splicing result of multiple target keyframes, and crop at least one detected page region to obtain at least one single page.

[0076] In this embodiment of the application, a page segmentation model (such as LayoutParser / DocBank / YOLOv8-Layout) can be used to detect the page regions in the splicing result of multiple target keyframes, and each page region can be cropped into a single page to obtain at least one single page.

[0077] The above achieves the conversion of physical content into digital images, breaking through the limitations of existing cross-platform content copying solutions, such as the need for a network, distance restrictions, complex operation steps, and the need to install the same application on multiple devices. It provides high-quality data input for subsequent multimodal content recognition and intelligent copying.

[0078] S102, perform multimodal recognition on each single page to obtain the structured data of each single page and output it. The structured data includes the multimodal data within the multimodal region of the single page.

[0079] In this embodiment, by integrating visual technology, deep learning, and multimodal large models, accurate recognition, classification, and storage of multimodal content such as text, QR codes, tables, and graphics can be achieved, thereby completing the processing and output of images into structured data. The multimodal data in each single page of structured data is stored according to the position of the multimodal regions (from top to bottom, from left to right).

[0080] Specifically, refined object detection and instance segmentation are performed using the deep learning-based document image analysis toolkit LayoutParser, the general document pre-trained model DocBank / PubLayNet, and the extended YOLOv8-seg instance segmentation model (supporting text, QR codes, tables, and graphics). Each single page is divided into different semantic regions, and each semantic region is treated as a multimodal region. The edge coordinates of each multimodal region are recorded and labeled with category tags (such as text, QR codes, tables, and graphics).

[0081] 1) The multimodal region is a text region, and the multimodal data is text data. A deep learning OCR model based on CRNN (Convolutional Recurrent Neural Network) and CTC (Connection Temporal Classification) is adopted to optimize the character set and feature extraction for multilingual scenarios, improve the mixed recognition of multilingual languages ​​and multi-directional text recognition (tilted text, mirrored text, reversed text), thereby achieving high-precision text data recognition and outputting the edge coordinates of each word / line, the content of the text data, and the recognition confidence.

[0082] 2) The multimodal region is the QR code region, and the multimodal data is the QR code data. The core decoding algorithm of the ZXing open-source library is used to identify the QR code region and QR code data, supporting mainstream matrix code formats such as QR code, Data Matrix, and Aztec Code. This embodiment introduces a distortion correction function based on perspective transformation. The QR code region is located through edge detection and quadrilateral fitting algorithms, and then homography matrix transformation is used to repair geometric distortions such as tilt and curvature, improving the error tolerance rate. Finally, the edge coordinates, text content, and recognition confidence score of the QR code are output.

[0083] 3) The multimodal region is a table region, and the multimodal data is table data. Based on a hybrid model (CNN feature extraction + Transformer global association analysis), the Table Master table recognition model performs row and column recognition, nested table recognition, and cross-row list table recognition on the table region, and restores the row and column structure. The final output includes a list of cell-level locations, a multidimensional array of the content of each cell, and the recognition confidence score.

[0084] 4) The multimodal region is a graphical region, and the multimodal data is graphical data. A finely tuned ImageNet pre-trained model (such as ResNet, ViT, etc.) is used to identify graphical regions and categories (logos, product design diagrams, business process diagrams, maps, faces, etc.). A multimodal large model (VLM) is then used to generate descriptions of the graphical regions (e.g., "This is a brand logo with red text and black text"). The graphical data is then converted to base64 format, and the final output includes the category, description, base64 formatted graphical data, and recognition confidence score of the graphical region.

[0085] Additionally, for base64 formatted image data, it can be stored in a local cache first, with only the cache path information stored in memory, thus reducing memory pressure.

[0086] S103, determine at least one target multimodal data that is selected under a single page, and copy the target multimodal data to the target editing location.

[0087] In this embodiment, after outputting the structured data for each single page, the user can select a target multimodal region within one or more single pages, thereby determining the target multimodal data to be copied. Furthermore, the user can select the target editing location to paste (such as a browser, chat window, Word document, etc.) and copy the target multimodal data to that location.

[0088] In one possible implementation, intelligent selection and manual selection can be combined to support users in selecting target multimodal data, improving the accuracy and efficiency of content selection for copying. See also Figure 4 , Figure 4 This is another part of the flowchart illustrating a cross-end copying method provided in an embodiment of this application. For example... Figure 4 As shown in the embodiment of this application, a cross-platform copying method is provided, wherein step S103, "determining at least one target multimodal data selected under a single page", may include steps S401 to S402, which are described in detail below.

[0089] S401, responding to the selection operation of the target multimodal region under the target single page, highlights the target multimodal region and uses the multimodal data within the target multimodal region as the target multimodal data.

[0090] In this embodiment, the user can sequentially select a target page to be copied from at least one single page. Clicking the target page triggers automatic selection of the target multimodal region corresponding to the clicked location, highlighting the target multimodal region. For example, the target multimodal region is highlighted with a blue border and overlaid with transparent fill, while simultaneously triggering vibration feedback to ensure the user clearly perceives the selected area. Double-clicking an already highlighted area cancels the selection and highlighting of that area. Therefore, the multimodal data within the target multimodal region can be used as the target multimodal data to be copied.

[0091] S402, when the target multimodal region is a table region or a text region, respond to the selection operation on the target multimodal data and filter the selected data from the target multimodal data.

[0092] In this embodiment, if the target multimodal region is a table region, in addition to supporting the above-mentioned single-click selection of all, it also supports a "long-press and drag" operation to select a portion of the data in each cell region, so as to filter the selected data from the full data. Furthermore, if the target multimodal region is a text region, in addition to supporting the above-mentioned single-click selection of all, it also supports a "long-press and drag" operation to select a portion of the data, so as to filter the selected data from the full data.

[0093] In one possible implementation, after the user selects the content and clicks "Confirm Copy," the selected multimodal data can be stored hierarchically. A smart menu can be generated to automatically filter data for relevant functions (such as "Paste Text Only," "Insert Table Only," "Extract All QR Codes," etc.) and paste it orderly into the target editing location, providing operational convenience and high performance efficiency. See also Figure 5 , Figure 5 This is another part of the flowchart illustrating a cross-end copying method provided in an embodiment of this application. For example... Figure 5 As shown in the embodiment of this application, a cross-end copying method is provided, wherein step S103, "copying the target multimodal data to the target editing location", may include steps S501 to S503, which are described in detail below.

[0094] S501, hierarchical storage of target multimodal data to obtain clipboard data.

[0095] In this embodiment, the hierarchical storage design balances structural clarity, query efficiency, and semantic scalability. When storing target multimodal data hierarchically, it first categorizes the data according to the data type of its respective target multimodal region (e.g., text, QR code, table, and graphics). Then, data of the same data type is sorted and stored according to the page number of its respective page. Finally, data within the same page is sorted and stored according to its spatial location (from top to bottom, from left to right). After hierarchical storage of the target multimodal data, the structured data obtained is used as clipboard data.

[0096] In one possible implementation, to avoid long-term occupation of system resources, this application embodiment can set the default lifespan of clipboard data based on an automatic expiration mechanism using timestamps. When this lifespan expires or a user performs a new copy operation, historical cached data can be automatically cleared. This optimizes memory usage while ensuring data availability, providing users with an efficient and convenient copy operation that is "copy once, available worldwide." In this regard, this application embodiment provides a cross-platform copying method, wherein step S502, "generating a function menu based on the clipboard data and the format types supported by the target editing location," can take the following steps:

[0097] Obtain the lifecycle of the clipboard data, and generate a function menu within the lifecycle based on the clipboard data and the format types supported by the target editing location.

[0098] In this embodiment, the lifespan of clipboard data can be obtained according to user configuration. Assuming the lifespan is set to 24 hours, if the user clicks on the target editing location within 24 hours, a corresponding function menu can be generated.

[0099] S502 generates a function menu based on clipboard data and the format types supported by the target editing location.

[0100] In this embodiment, different target editing locations support different format types. For example, a browser address bar only supports text links, a chat box supports both text links and image formats, and a Word document supports all data formats. Therefore, menu items can be automatically generated based on the data type in the clipboard data and the format types supported by the target editing location. For instance, the browser address bar's function menu will display three primary functions: "Paste All," "Paste Text," and "Paste QR Code." If there are multiple data entries under the QR code type, secondary functions such as "Paste All," "Paste https: / / example1.com," and "Paste https: / / example2.com" will be displayed for the user to choose from.

[0101] S503 responds to the selection operation of the function menu, determines the target function selected in the function menu, and copies the target data that matches the target function from the clipboard data to the target editing location.

[0102] In this embodiment, the user can select any function from the function menu. The selected function is then used to determine the target function in the menu, thereby copying the target data matching the target function from the clipboard data to the target editing location. For example, if the user selects the secondary function "Paste QR Code" - "Paste https: / / example1.com", the corresponding QR code content from the clipboard data will be pasted to the target editing location.

[0103] Based on the above description, the cross-platform copying method provided in this application, through a technical architecture of "photographic input - pagination processing - multimodal recognition - hierarchical storage - intelligent pasting," achieves direct copying from physical space to digital space, ensuring real-time performance and compatibility of cross-platform data transmission while maintaining ease of operation. This application has the following advantages:

[0104] 1) Breakthrough in environmental adaptability: By acquiring and recognizing images, it breaks free from network dependence and device pairing limitations, overcomes the environmental limitations of traditional cloud synchronization, and supports cross-device replication capabilities.

[0105] 2) Breakthrough in copying content types: Through image acquisition and recognition, it supports multiple types of information such as screen content and paper documents, and the scenario coverage is greatly expanded compared with in-application tool solutions.

[0106] 3) Expanded multimodal data copying capabilities: Supports cross-platform copying of various information types such as text, QR codes, links, images, and tables, covering various mainstream scenarios, and has greatly improved compatibility compared to single-function tools.

[0107] 4) Intelligent selection and intelligent menu make copying convenient and efficient: Through multimodal region recognition and intelligent selection fusion technology, layered storage and intelligent paste menu improve the user experience of accurately selecting and copying complex multimodal content.

[0108] The above describes a cross-end copying method provided by the embodiments of this application. The following describes the apparatus for performing the above cross-end copying method.

[0109] See Figure 6 , Figure 6 This is a schematic diagram of a cross-end copying device provided in an embodiment of this application. Figure 6 As shown in the figure, an embodiment of this application provides a cross-end copying device, comprising:

[0110] The pagination processing module 601 is used to capture a copied image through a camera and perform pagination processing on the copied image to obtain at least one single page.

[0111] The multimodal recognition module 602 is used to perform multimodal recognition on each single page to obtain the structured data of each single page and output it. The structured data includes the multimodal data within the multimodal region of the single page.

[0112] The data copy module 603 is used to determine at least one target multimodal data that is selected in a single page and copy the target multimodal data to the target editing position.

[0113] In one possible implementation, a paging processing module 601 for paging the copied screen to obtain at least one single page is specifically used for:

[0114] When the copied image is a single image, the copied image is denoised and cropped to obtain the target image; page detection is performed on the target image, and at least one detected page region is cropped to obtain at least one single page; at least one single page is deduplicated based on perceptual hashing, feature matching, and OCR recognition.

[0115] In one possible implementation, a paging processing module 601 for paging the copied screen to obtain at least one single page is specifically used for:

[0116] When the copied image is a video, multiple keyframes are extracted from the copied image through scene detection and layout detection; multiple keyframes are deduplicated based on the inter-frame difference method; multiple target keyframes are determined by clustering the deduplicated results of multiple keyframes, and multiple target keyframes are stitched together. The multiple target keyframes belong to different categories, and each target keyframe is the keyframe with the highest clarity in its category; page detection is performed on the stitched result of multiple target keyframes, and at least one detected page area is cropped to obtain at least one single page.

[0117] In one possible implementation, the data copying module 603 for determining at least one selected target multimodal data within a single page is specifically used for:

[0118] In response to a selection operation on a target multimodal region within a target single page, the target multimodal region is highlighted, and the multimodal data within the target multimodal region is used as the target multimodal data. If the target multimodal region is a table area or a text area, in response to a selection operation on the target multimodal data, the selected data is filtered from the target multimodal data.

[0119] In one possible implementation, the data copying module 603, used to copy the target multimodal data to the target editing location, is specifically used for:

[0120] The target multimodal data is stored in layers to obtain clipboard data; a function menu is generated based on the clipboard data and the format types supported by the target editing location; in response to the selection operation of the function menu, the target function selected in the function menu is determined, and the target data matching the target function in the clipboard data is copied to the target editing location.

[0121] In one possible implementation, the data copy module 603, used to generate the function menu based on the clipboard data and the format type supported by the target editing location, is specifically used for:

[0122] Obtain the lifecycle of the clipboard data, and generate a function menu within the lifecycle based on the clipboard data and the format types supported by the target editing location.

[0123] It should be noted that the detailed functions of each module in the embodiments of this application can be found in the corresponding disclosure of the cross-end copying method embodiments described above, and will not be repeated here.

[0124] This application also provides an electronic device in its embodiments. See also... Figure 7 , Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device in this embodiment may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0125] like Figure 7 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. When the electronic device is powered on, the RAM 703 also stores various programs and data required for the operation of the electronic device. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0126] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, memory cards, hard drives, etc.; and communication devices 709. Communication device 709 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.

[0127] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the cross-end copying methods provided in this application.

[0128] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the cross-end copying methods provided in this application.

[0129] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0130] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0131] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.

[0132] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A cross-end copying method, characterized in that, The cross-end copy method includes: The copied image is captured by a camera, and the copied image is paginated to obtain at least one single page; Multimodal recognition is performed on each single page to obtain the structured data of each single page and output it. The structured data includes the multimodal data within the multimodal region of the single page. The system identifies the target multimodal data selected within at least one single page and stores it hierarchically to obtain clipboard data. During hierarchical storage, the target multimodal data is categorized according to the data type of its respective target multimodal region, and data of the same data type is sorted and stored according to the page number of its respective single page. Data within the same single page is sorted and stored according to its spatial location. A function menu is generated based on the clipboard data and the format types supported by the target editing location. In response to a selection operation on the function menu, the system identifies the selected target function and copies the target data from the clipboard data that matches the target function to the target editing location, thus enabling direct copying from physical space to digital space without a network connection.

2. The cross-end copying method according to claim 1, characterized in that, The step of paginating the copied screen to obtain at least one single page includes: If the copied image is a single image, the copied image is denoised and cropped to obtain the target image; The target screen is subjected to page detection, and at least one detected page area is cropped to obtain at least one single page; The at least one single page is deduplicated based on perceptual hashing, feature matching, and OCR recognition.

3. The cross-end copying method according to claim 1, characterized in that, The step of paginating the copied screen to obtain at least one single page includes: When the copied image is a video, multiple keyframes are extracted from the copied image through scene detection and layout detection. The multiple keyframes are deduplicated based on the inter-frame difference method; Multiple target keyframes are determined by clustering the deduplication results of the multiple keyframes, and the multiple target keyframes are stitched together. The multiple target keyframes belong to different categories, and each target keyframe is the keyframe with the highest clarity in its category. Page detection is performed on the splicing result of the multiple target keyframes, and at least one detected page region is cropped to obtain at least one single page.

4. The cross-end copying method according to claim 1, characterized in that, The determination of the target multimodal data selected in the at least one single page includes: In response to a selection operation of a target multimodal region under a target single page, the target multimodal region is highlighted, and the multimodal data within the target multimodal region is used as the target multimodal data; When the target multimodal region is a table region or a text region, in response to the selection operation on the target multimodal data, the selected data is filtered from the target multimodal data.

5. The cross-end copying method according to claim 1, characterized in that, The step of generating a function menu based on the clipboard data and the format types supported by the target editing position includes: Obtain the lifespan of the clipboard data, and generate the function menu within the lifespan based on the clipboard data and the format type supported by the target editing position.

6. A cross-end copying device, characterized in that, The cross-end copying device includes: The pagination processing module is used to capture a copied image through a camera and perform pagination processing on the copied image to obtain at least one single page; A multimodal recognition module is used to perform multimodal recognition on each single page to obtain and output the structured data of each single page. The structured data includes multimodal data within the multimodal region of the single page. The data copy module is used to determine the target multimodal data selected in the frame under the at least one single page, and copy the target multimodal data to the target editing position; Specifically, the data copying module, which copies the target multimodal data to the target editing location, is used for: The target multimodal data is stored in layers to obtain clipboard data. During this layered storage, the target multimodal data is categorized according to the data type of its respective target multimodal region. Target multimodal data of the same data type are sorted and stored according to the page number of their respective single page. Target multimodal data within the same single page are sorted and stored according to spatial location information. A function menu is generated based on the clipboard data and the format types supported by the target editing location. In response to a selection operation on the function menu, the selected target function is determined, and the target data matching the target function in the clipboard data is copied to the target editing location, thus enabling direct copying from physical space to digital space without a network connection.

7. A computer program product, characterized in that, Includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the cross-end copy method as described in any one of claims 1 to 5.

8. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the cross-end copy method as described in any one of claims 1 to 5.

9. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the cross-end copy method as described in any one of claims 1 to 5.