Data processing method and equipment
By analyzing user selection actions using a target processing model, selection results that match user intent are generated, solving the problem of inaccuracy in manual selection, achieving more efficient and accurate content selection, and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-17
AI Technical Summary
During user interaction with electronic devices, manual selection is difficult to accurately select the content area the user expects, resulting in a discrepancy between the selection range and the actual intent, leading to over- or under-selection, which affects the accuracy of subsequent processing and user experience.
By analyzing the user's selection trajectory using a target processing model, target analysis results are generated, and a second region matching the user's intent is identified, enabling automatic correction and precise positioning of the user's selected content. This model includes visual feature extraction, OCR, layout/structure analysis, and a cross-modal fusion and intent inference layer. It combines contextual information to analyze semantic and spatial relationships, generating more accurate selection results.
It improves the accuracy and completeness of selection results, reduces the burden of repeated adjustments for users, and enhances user interaction experience and processing efficiency.
Smart Images

Figure CN121879644A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a data processing method and apparatus. Background Technology
[0002] During the interaction between a user and an electronic device, the user can select certain content on the display interface and then further process the selected content. However, due to the limitations of manual operation, it is often difficult to accurately select the content area that the user expects, resulting in a deviation between the final selection range and the user's true intention. Summary of the Invention
[0003] In view of this, this disclosure provides a data processing method.
[0004] According to a first aspect of this disclosure, a data processing method is provided, comprising: obtaining first display content of a display interface; in response to a user's selection operation on the first display content, determining a first region corresponding to the selection operation; obtaining the first region based on a first trajectory formed by the selection operation; analyzing and processing a second display content corresponding to the first region using a target processing model to generate a target analysis result; determining a second region based on the target analysis result; the second region including at least a portion of the first region.
[0005] A second aspect of this disclosure provides an electronic device, comprising: at least one processor and a target application running on the processor, the target application being capable of independently executing or invoking at least one artificial intelligence model to perform the following operations: obtaining first display content of a display interface; determining a first region corresponding to the selection operation in response to a user's selection operation on the first display content; the first region being obtained based on a first trajectory formed by the selection operation; analyzing and processing a second display content corresponding to the first region using a target processing model to generate a target analysis result; determining a second region based on the target analysis result; the second region including at least a portion of the first region.
[0006] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0007] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0008] Figure 1 A schematic diagram illustrating a data processing method in the related art is shown.
[0009] Figure 2A flowchart illustrating a data processing method according to an embodiment of the present disclosure is shown schematically.
[0010] Figure 3 One of the schematic diagrams of a data processing method according to an embodiment of the present disclosure is shown;
[0011] Figure 4 A second schematic diagram of a data processing method according to an embodiment of the present disclosure is shown.
[0012] Figure 5 A schematic diagram of a data processing method according to an embodiment of the present disclosure is shown in Figure 3.
[0013] Figure 6 The diagram illustrates a fourth schematic of a data processing method according to an embodiment of the present disclosure. Detailed Implementation
[0014] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0015] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0016] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0017] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0018] This disclosure provides a data processing method. Before introducing the technical solutions provided by this disclosure, the relevant technologies involved in this disclosure will be described first.
[0019] During user interaction with electronic devices, users can select certain content on the display interface and then further process the selected content. However, due to the limitations of manual operation, it is often difficult to accurately select the content area desired by the user, resulting in a deviation between the final selection range and the user's true intention. This is mainly manifested in the following ways:
[0020] (1) Multiple selection: When selecting target content, users tend to include non-target parts around the target, such as blank areas, adjacent text, other paragraphs, image backgrounds or unrelated graphic elements, resulting in redundant selection range.
[0021] (2) Incomplete selection: Users may miss part of the target content, such as the last character of a word, several lines in a paragraph, part of the edge of a graphic or the outline of an object, resulting in incomplete selection.
[0022] In one example, such as Figure 1 As shown, in the text interface, a user wants to select "The CEO of Company Ais BAC". However, due to the precision limitations of finger selection or existing intelligent agent selection, the actual selection area may contain extra blank spaces or adjacent words (multiple circles), or omit part of the name (incomplete selection). This kind of inaccurate selection area will directly affect the accuracy of subsequent intelligent processing, and cause the user to need to adjust the selection area multiple times or even redo the operation, thereby reducing user experience and work efficiency.
[0023] The following will be through Figures 2-6 The data processing method of the embodiments of this disclosure will be described in detail.
[0024] Figure 2 A flowchart illustrating a data processing method according to an embodiment of the present disclosure is shown schematically.
[0025] like Figure 2 As shown, the data processing method of this embodiment includes operations S210 to S240.
[0026] During operation S210, the first display content of the display interface is obtained.
[0027] In operation S220, in response to the user's selection operation on the first display content, the first area corresponding to the selection operation is determined; the first area is obtained based on the first trajectory formed by the selection operation.
[0028] In operation S230, the target processing model is used to analyze and process the second display content corresponding to the first area, and the target analysis result is generated.
[0029] In operation S240, based on the target analysis results, a second region is determined; the second region includes at least a portion of the first region.
[0030] For example, the display interface may be the display area of an electronic device with a display screen.
[0031] The first displayed content can be the visible content presented on the current display interface. For example, the first displayed content can be an image, text, chart, webpage, video frame, etc. Similarly, the first displayed content can also include the visible content of the original document and supplementary content that complements the visible content. For example, if some semantics in the current document are incomplete, they can be analyzed and supplemented based on the original content to make the semantically incomplete parts of the current document complete, and then the supplemented visible content and the supplementary content can be displayed together on the display interface. The first displayed content can also be an image captured from the current display interface. For example, if the current display interface shows a document to read, a screenshot of the current reading document page can be used as the first displayed content. Selection operations can be user interactions such as mouse dragging, touchscreen drawing, gestures, or selection tools. Selection operations include, but are not limited to: long press and slide, drag selection, free curve selection, smudge selection, lasso selection, etc.
[0032] The display interface provides the visual and operational environment for user selection and serves as an instantaneous display layer for the intelligent selection results. The system responds immediately when the user performs mouse drags or gestures on this interface. For example, a high-performance graphics API (Application Programming Interface) can be used, such as the graphics interfaces corresponding to different operating systems, like Windows' Direct2D (Direct2D Graphics API), to achieve GPU (Graphics Processing Unit) accelerated rendering. This ensures that when the user drags the mouse or makes movement gestures, the intelligent selection can smoothly follow, predict, and "snap" onto the target content at a very high frame rate (e.g., above 60fps), avoiding any stuttering or visual lag. The selection area can be represented as a highlight mask, a precise outline, or a fill color with transparency.
[0033] It can capture pixel-level image data of the user-selected area with extremely low latency, while simultaneously recording the user's original precise selection trajectory (pixel coordinate sequence) and geometric information (such as start point, end point, width, and height). For paper documents, it can acquire high-resolution video streams in real time via a camera, and may have built-in lightweight image correction algorithms (such as distortion correction and perspective transformation). It utilizes the underlying APIs of the operating system, using the DXGI Desktop Duplication API on Windows and the Core Graphics Window List CreateImage function on Mac. These APIs allow direct access to framebuffer data, minimizing data copy overhead. The captured image data is directly passed to the underlying input layer of the VLM (Video Language Model) in a zero-copy or efficient memory-mapped manner, avoiding unnecessary CPU (Central Processing Unit)-GPU data transfer bottlenecks. Capturing the x, y coordinate sequence of mouse or touch events provides crucial clues for understanding the user's intent.
[0034] The first trajectory can be the path of input points generated when the user performs a selection operation. For example, if the selection operation is a drag-and-drop selection: the first trajectory can be the trajectory from the pressed point to the released point; if the selection operation is a free curve: the first trajectory can be a set of continuous curve points.
[0035] The first region can be a geometric area within the first displayed content defined by the first trajectory of the user's selection operation on the first displayed content. For example, if the selection operation is a drag-and-drop selection, the user presses and drags on the document to form a rectangle, and the first region is determined by the coordinates of the top-left and bottom-right corners of the rectangle. If the selection operation is a free-form curve selection, the user draws a circle around a chart with a stylus, and the area inside the curve is defined as the first region. If the selection operation is a smudge selection, the user smudges several times on the image with their finger, and the set of pixels covered by the smudge trajectory (which can be expanded by a certain radius) is defined as the first region.
[0036] The second displayed content can be content extracted from the first displayed content and related to the first area. For example, the second displayed content can also be a portion of the first displayed content corresponding to the first area; for instance, if the user selects the first and second cells in Table 1 on the third page of the displayed document, the second displayed content can be the first and second cells. If the user selects the head of a puppy in an image displayed on the interface, the second displayed content can be the puppy's head. If the user slides below the words "birds in the sky" in the third line of the third page of the displayed document, the second displayed content can be: birds in the sky. The second displayed content can also be a portion of the first displayed content larger than the first area; for example, if the user selects a portion of the cells in Table 1 on the third page of the displayed document, the second displayed content can be all the cells in Table 1, or it can be the cells corresponding to Table 1 and the first paragraph of text below Table 1.
[0037] The target processing model can be a traditional algorithm model, a machine learning model, or a deep learning model, including but not limited to: multimodal large models, chart understanding models, and text classification models. The analysis of the second display content corresponding to the first area by the target processing model includes, but is not limited to, identifying the object categories within the area, recognizing and understanding text semantics, and analyzing the page layout structure.
[0038] The second displayed content is input into the target processing model. The target processing model performs semantic analysis and / or layout analysis on the content with the second displayed content and outputs the target analysis results. The target analysis results can be the result data output by the target processing model after analyzing and processing the second displayed content corresponding to the first region. The target analysis results are used to characterize the target processing model's understanding, recognition, or inference conclusions about the content the user expects to select. For example, the target analysis results can be text or image content, or a trajectory corresponding to the text or image content.
[0039] The second region can be the area containing the displayed content corresponding to the target analysis result. The second region overlaps with the first region at least partially; that is, the displayed content in the second region includes at least a portion of the displayed content in the first region. For example, the second region can be the same as the first region; for instance, if a user selects Table 1 (the second displayed content) on the third page of a document displayed on the interface, inputs the second displayed content into the large language model, and the large language model outputs Table 1 as the target analysis result, the area corresponding to Table 1 is the second region, then the first and second regions are the same. The second region can also be a larger region that includes the entire first region. For example, if a user selects a portion of Table 1 (the second displayed content) on the third page of a document displayed on the interface, inputs the second displayed content and the third page of the document into the large language model, and the large language model outputs Table 1 as the target analysis result, the area corresponding to Table 1 and the first paragraph below Table 1 is the second region, then the first region is smaller than and the second region is the same. The second region can also be a smaller region that includes a portion of the first region. For example, if a user selects Table 1 (the second displayed content) on the third page of the document displayed on the screen, and inputs the third page of the second displayed content into the large language model, the target analysis result output by the large language model is the title of Table 1. The area corresponding to the title of Table 1 is the second area, then the first area is larger than the second area.
[0040] In one example, a user wants to select the phrase "birds fly in the sky" in the second line of the fourth page of a displayed document. The user uses a stylus to slide across the text below the second line of the document, with the starting point of the slide being "wind" and the ending point being "sky." The first area covered by this slide is the text "wind, birds fly in the sky." The fourth page of the document with the user's slide is input into a large language model, resulting in the phrase "birds fly in the sky." The area containing "birds fly in the sky" is then designated as the second area, representing the user's final selection result, and is available for subsequent operations such as copying, annotation, and editing.
[0041] Understandably, the model analyzes the displayed content corresponding to the user's initial selection trajectory to generate target analysis results. Based on these results, a second region matching the user's intent is determined, enabling automatic correction and precise positioning of the user's selected content. This improves the accuracy of analysis and processing, as well as interaction efficiency. Furthermore, the second region at least partially overlaps with the first region, ensuring it still revolves around the user's selection intent. Therefore, without requiring additional fine-tuning by the user, it reduces interference from invalid regions, improves the accuracy and stability of subsequent recognition, and enhances the user experience and overall processing efficiency.
[0042] As described above, in operation S230, the target processing model is used to analyze and process the second display content corresponding to the first region to generate a target analysis result. In one possible implementation, this operation may further include the following steps: using the target processing model to analyze and process the first and second display content to obtain contextual information of the second display content; and using the contextual information to edit the second display content to obtain the target analysis result.
[0043] For example, the contextual information can be surrounding content information that is related to the second displayed content. The contextual information can originate from other parts of the first displayed content besides the second displayed content. For instance, if the second displayed content is a portion of text within a paragraph, the contextual information can include other text in that paragraph, content from adjacent paragraphs, page titles, etc. If the second displayed content is a region within an image, the contextual information can include other regions of that image, image annotations, image description text, etc.
[0044] When the target processing model analyzes and processes the first and second displayed content, it can comprehensively consider various factors such as the positional relationship, semantic relationship, and structural relationship of the second displayed content within the first displayed content, thereby extracting contextual information related to the second displayed content.
[0045] Editing can be an operation that expands, reduces, corrects, or reorganizes the second displayed content based on contextual information. For example, editing can include: expanding the boundaries of the second displayed content based on contextual information to include semantically related surrounding content; semantically completing the second displayed content based on contextual information to add missing semantic units; and filtering the second displayed content based on contextual information to remove irrelevant and redundant parts.
[0046] Specifically, the target processing model can be a local multimodal intelligent model (VLM), and the input layer of the target processing model can directly receive an image tensor with the first display content and the first trajectory. The target processing model should achieve end-to-end learning as much as possible, that is, a direct mapping from the original pixels to the final second region, reducing the concatenation of intermediate modules to reduce latency.
[0047] The target processing model includes: a visual feature extraction layer, an OCR (Optical Character Recognition) and text encoding layer, a layout / structure analysis layer, and a cross-modal fusion and intent inference layer.
[0048] The visual feature extraction layer can use a miniaturized Vision Transformer (ViT) as an image encoder, which can extract multi-scale, high-dimensional visual features from the input image in parallel, such as text regions, image content, boundary lines, icons and other visual elements.
[0049] VLM contains an OCR module. This is not a separate external OCR engine, but rather part of VLM. It works directly on visual features through a Text Recognition Head to output text sequences. These text sequences are then transformed into semantic embeddings using a lightweight text encoder, which can be distilled using Distil BERT (a distilled version of the Transformer-based bidirectional encoder representation), providing contextual information.
[0050] VLM can include a layout analysis sub-network. This sub-network identifies logical structural units on the page (such as paragraphs, headings, list items, table cells, code blocks, etc.) based on visual features and text embeddings, and generates structural features.
[0051] Cross-modal fusion and intent inference layer: This layer is the core decision-making unit of VLM. It deeply fuses visual, contextual, and structural features through a cross-modal attention mechanism, while also incorporating geometric information from the user's original selection (such as center point, aspect ratio, area, and trajectory shape). VLM understands the user's intent at this layer and ultimately inputs the target analysis results. For example, it determines whether the user wants to select "a complete word" rather than a part, "a paragraph" rather than several lines, or "a specific object in an image" rather than the entire image.
[0052] In one example, a user wants to select the complete description of "2023 Q2 sales growth of 15%" from page 5 of a research report displayed on the screen. The user drags a selection box on the sentence with the mouse, but due to imprecise operation, the selection starts at "quarter" and ends at "growth of 15", with the first area covering the text "quarterly sales growth of 15%". The fifth page of the research report (the first displayed content) and the corresponding "quarterly sales growth of 15%" (the second displayed content) are input into the target processing model. The target processing model analyzes the fifth page and the second displayed content, identifying the sentence structure of the second displayed content. It finds that "quarter" is preceded by the time qualifier "2023 Q2", and "15" is immediately followed by a percentage sign "%", these surrounding elements constitute the contextual information of the second displayed content. Based on this contextual information, the target processing model determines that the user's intent is to select a complete data description statement. Therefore, it edits the second displayed content, adding "2023 Second" and the "%" to obtain the complete "2023 Second Quarter Sales Growth of 15%" as the target analysis result. Subsequently, the area containing "2023 Second Quarter Sales Growth of 15%" is designated as the second region, serving as the user's final selection result.
[0053] Understandably, by identifying contextual information in the first displayed content and editing the second displayed content based on that contextual information, it's possible to more accurately infer the user's true selection intent. This allows for the automatic completion of missing semantic units or the removal of redundant content, achieving intelligent optimization of the user's initial selection result. Furthermore, inputting both the first and second displayed contents into the target processing model for analysis fully utilizes the overall information of the page to aid in judgment. Therefore, while maintaining processing efficiency, it can significantly improve the accuracy and completeness of the selection result, reduce the burden of repeated adjustments for the user, and improve the overall interactive experience.
[0054] Figure 3 One of the schematic diagrams of a data processing method according to an embodiment of the present disclosure is shown.
[0055] As described above, the operation involves using a target processing model to analyze and process the first and second displayed content to obtain the contextual information of the second displayed content. In one possible implementation, this operation may further include the following steps: performing semantic analysis on the first and second displayed content, using the third displayed content within the first displayed content as the contextual information of the second displayed content; and ensuring that the semantic relevance between the third displayed content and the second displayed content is greater than a preset relevance threshold.
[0056] In another possible implementation, the operation may further include the following operations: performing layout analysis on the first display content and the second display content, using the fourth display content in the first display content as context information of the second display content; the fourth display content and the second display content satisfy a preset positional relationship.
[0057] For example, semantic analysis can be a process of understanding and judging the meaning, theme, and conceptual relationships of text, images, or combinations thereof. For instance, semantic analysis can identify that "sales growth" and "performance improvement" have similar semantics, "second quarter" and "Q2" represent the same time concept, and "product A" and "this product" have a referential relationship, etc.
[0058] The third displayed content can be a portion of the first displayed content that is semantically related to the second displayed content. The semantic relevance between the third displayed content and the second displayed content is greater than a preset relevance threshold (e.g., 90%). For example, if the second displayed content is "15% growth", the third displayed content could be "sales revenue in the second quarter of 2023" from the same paragraph, because these two parts together constitute a complete data description, and their semantic relevance is higher than the preset relevance threshold (90%).
[0059] A preset relevance threshold can be a numerical standard used to determine whether the semantic relevance is sufficient. It should be noted that the preset relevance threshold can be set and adjusted according to factors such as application scenario, content type, and user habits.
[0060] Layout analysis is the process of identifying and understanding the spatial location, structural hierarchy, and formatting of displayed content. It can identify structural units on a page, such as paragraphs, headings, lists, tables, images, headers, and footers, and determine the positional and hierarchical relationships between these units. For example, layout analysis can identify a paragraph of text located in the upper left corner of the page and belonging to the heading level; a table located in the middle of the page, containing three rows and five columns; and an image located below a paragraph of text, indicating a descriptive relationship with that text.
[0061] The fourth displayed content can be a portion of the first displayed content that has a specific positional relationship with the second displayed content. Preset positional relationships can include, but are not limited to: adjacent relationships, containment relationships, alignment relationships, same-line relationships, same-column relationships, same-paragraph relationships, and same-section relationships. For example, if the second displayed content is a cell in a table, the fourth displayed content can be other cells in the same row as that cell, because they satisfy the same-line relationship. If the second displayed content is part of the text in a paragraph, the fourth displayed content can be other text in that paragraph, because they satisfy the same-paragraph relationship; the fourth displayed content can also be the paragraph immediately preceding or following that paragraph, because they satisfy the adjacent relationship.
[0062] In other embodiments, semantic analysis and layout analysis can be combined to determine relevant display content from the first display content as context information for the second display content.
[0063] In one example, if a user wants to select Table 1 to display on the interface, the system first precisely locates the user-specified Table 1. Then, by analyzing its positional relationships, the system extracts the preceding and following paragraphs adjacent to Table 1 as the layout context. Simultaneously, by analyzing the semantic relationships of the text, the system extracts the core paragraphs that specifically describe or explain the content of Table 1 as the semantic context. Finally, the layout context and semantic context are integrated to form the contextual information for the second displayed content.
[0064] Understandably, by performing semantic analysis on the first and second displayed content, semantically related content can be identified as contextual information, thereby understanding the user's intent based on the meaning of the content and enabling intelligent completion of semantically incomplete selections. Simultaneously, by performing layout analysis on the first and second displayed content, location-related content can be identified as contextual information, thereby understanding the user's intent based on spatial structure and achieving precise positioning of structured content.
[0065] As described above, the operation involves editing the second displayed content using contextual information to obtain the target analysis result. In one possible implementation, this operation may further include: editing the second displayed content based on at least one fifth displayed content from the contextual information to obtain the target analysis result; the semantic relevance between at least one fifth displayed content and at least a portion of the content in the second displayed content satisfies a preset condition.
[0066] In another possible implementation, the operation may further include: editing the second display content based on a preset identifier in the context information to obtain the target analysis result.
[0067] In other embodiments, the second display content can be edited based on at least one fifth display content in the context information and a preset identifier in the context information to obtain the target analysis result.
[0068] For example, the fifth displayed content can be a content unit in the context information that has a specific semantic relationship with the second displayed content. The fifth displayed content can be a content fragment with clear semantics, such as keywords, titles, terms, entity names, time statements, numerical statements, etc. For example, if the second displayed content is "15% growth", the fifth displayed content can be the keyword "sales revenue" in the context information, or the time statement "second quarter of 2023", or the title "performance report".
[0069] If at least one fifth piece of displayed content satisfies a preset condition in semantic relevance with at least a portion of the second displayed content, it indicates that the fifth and second displayed content are semantically substitutable, complementary, or continuous. The preset condition may include: a semantic relevance higher than a preset relevance threshold, indicating that the fifth and second displayed content are semantically close and can be used for completion or replacement; or a semantic similarity lower than a preset relevance threshold, indicating that the fifth and second displayed content have significant semantic differences, potentially indicating user error and requiring further confirmation.
[0070] For example, if the second displayed content is "Product A Performance", and the context information contains two fifth displayed contents, "Product B Function" and "Product A Performance Parameter", the former has a low semantic similarity to the second displayed content, while the latter has a high semantic similarity to the second displayed content. If the preset conditions are met, then the latter will be selected for editing.
[0071] In one scenario, if all identified fifth-display content has a low semantic similarity to the second-display content (i.e., all semantic similarities are below a preset similarity threshold), it indicates a significant error in the user's selection or a semantic conflict between the contextual information and the second-display content. In this case, a prompt can be generated, requiring the user to provide additional confirmation. For example, if the second-display content is "sales growth," but the contextual information contains semantically opposite terms such as "cost reduction" or "profit reduction," or contains typos like "sales growth," then all fifth-display content will have a low semantic similarity to the second-display content. The system can highlight these ambiguous or erroneous contents and prompt the user to reselect or confirm their selection.
[0072] Preset identifiers can be characters or symbols that have a specific separating or marking function in the displayed content. Preset identifiers can include, but are not limited to: punctuation marks such as spaces, periods, commas, semicolons, colons, question marks, exclamation marks, line breaks, tabs, quotation marks, parentheses, and book titles, as well as structural identifiers such as paragraph marks, list symbols, and table boundaries.
[0073] Editing the second displayed content based on preset identifiers allows for the determination of its logical boundaries, enabling expansion or truncation. For example, if the second displayed content is "quarterly sales growth," the selection trajectory actually starts from the word "quarter." However, the target processing model identifies preset identifiers "." (period) before and "," (comma) after the second displayed content in the context, such as "2023 second quarter sales growth 15%." Therefore, it can determine that the complete semantic unit should be "2023 second quarter sales growth 15%," and thus complete the second displayed content based on these preset identifiers.
[0074] In one example, a user selects content in a project progress report using discontinuous smearing actions (covering three fragments: "mileage", "prototype development", and "March 2024"). The target processing model processes the input content by combining semantic relevance analysis and pre-defined identifier recognition. First, it analyzes the contextual information, finding that the fragment mileage has a high correlation with the complete semantic unit key milestone in the context, prototype development is closely related to prototype development completion, and March 2024 is part of March 15, 2024. Simultaneously, it identifies the colon ":" as a separator between labels and content, and the period and newline character as sentence boundary identifiers. Based on semantic relevance and pre-defined identifiers, the discrete fragments are completed into a complete sentence: Key Milestone: Prototype development completion time is March 15, 2024, which is then identified as the target analysis result. This process, relying solely on semantic relevance and pre-defined identifiers, achieves the recovery from partial, discontinuous input to a complete and accurate semantic unit.
[0075] In another example, the user selects "Second Quarter Sales" and the context information contains the corresponding "Second Quarter Sales Amount". At this point, it can be determined that "Second Quarter Sales Amount" is closely related to "Second Quarter Sales". By expanding "Second Quarter Sales", we can obtain the target analysis result of "Second Quarter Sales Amount".
[0076] Understandably, by editing the second displayed content based on the fifth displayed content in the context information, semantic similarity matching can be used to automatically identify and complete the complete semantic unit of the user's intended selection, thereby correcting content gaps caused by discontinuous or incomplete selection operations and improving the semantic integrity of the selection result. Simultaneously, by editing the second displayed content based on preset identifiers in the context information, the structural separation features of the content can be used to accurately define semantic boundaries, automatically remove redundant parts or expand to complete units, thereby correcting boundary deviations caused by inaccurate selection ranges and improving the structural integrity of the selection result.
[0077] As described above, the operation involves editing the second displayed content using context information to obtain the target analysis result. In one possible implementation, this operation may further include the operation that, if the second displayed content includes image content, at least one image element in the image content is edited based on context information to obtain the target analysis result.
[0078] For example, the image content may be visual elements such as pictures, charts, icons, photos, and diagrams contained in the second display content.
[0079] Editing at least one image element in the image content based on contextual information may include, but is not limited to: retaining a specific image element, deleting redundant image elements, expanding the range of an image element, and adjusting the boundaries of an image element.
[0080] Figure 4 A second schematic diagram of a data processing method according to an embodiment of the present disclosure is shown.
[0081] In some embodiments, the operation includes: if the second displayed content includes image content, editing at least one image element in the image content based on context information to obtain a target analysis result. In one possible implementation, the operation may further include: if the context information includes text content, identifying the semantic description corresponding to at least one image element in the context information; deleting at least one image element that does not correspond to the semantic description to obtain a target analysis result.
[0082] For example, a semantic description can be a description of the meaning, attributes, or functions of an image element by the text content in the context information. For instance, if the second displayed content is a product display image, and the text content in the context information includes "This product's main body is made of metal," this text content constitutes a semantic description of the image element "metal body."
[0083] Deleting at least one of the image elements based on semantic description can be achieved by analyzing the objects mentioned or emphasized in the text content, identifying which image elements in the second display content are related to the semantic description, and removing irrelevant image elements from the second display content.
[0084] In one example, refer to Figure 4 If a user selects two puppies and a kitten in an image containing multiple animals, but the context indicates that the user is interested in a specific object: the two puppies, the target processing model can delete other irrelevant objects, such as the kitten, and retain only the target object, the two puppies, as the target analysis result. If a user selects a part of a chart, but the context indicates that the chart is a whole semantic unit, the target processing model can expand the selection range to include the entire chart.
[0085] Figure 5 The diagram illustrates a third schematic of a data processing method according to an embodiment of the present disclosure.
[0086] In another possible implementation, the operation may further include the following operations: if the context information includes image content, identifying a target image element in the context information; editing at least one image element in the image content corresponding to the second display content based on the target image element to obtain a target analysis result; the similarity and / or correlation between the target image element and at least one image element in the image content corresponding to the second display content is greater than a target threshold.
[0087] For example, the target image element may be an image element in the image content of the context information that is associated with the image content of the second display content.
[0088] Similarity can be defined as the degree of visual similarity between a target image element and an image element in the corresponding image content of the second display content, including but not limited to: color similarity, shape similarity, texture similarity, and size similarity. Association can be defined as the degree of semantic or spatial association between a target image element and an image element in the corresponding image content of the second display content, including but not limited to: positional association, action association, logical association, and temporal association. For example, positional association can be two image elements being symmetrical, adjacent, or in the same row or column in the page layout; action association can be two image elements representing different stages of the same operation process or a causal relationship; logical association can be two image elements belonging to the same classification system or hierarchical structure; and temporal association can be two image elements representing a chronological order.
[0089] The target threshold can be a preset value used to determine whether the similarity and / or relevance is high enough. For example, the target threshold can be a percentage of the similarity score (such as 80% or 90%), a numerical range of the relevance score (such as 0.7 to 1.0), or a weighted score of multi-dimensional comprehensive evaluation.
[0090] Editing at least one image element in the image content corresponding to the second display content based on the target image element can include, but is not limited to: adding the target image element, deleting image elements unrelated to the target image element, adjusting the boundaries of image elements to align with the target image element, and expanding the range of image elements to include other elements associated with the target image element. For example, if the target image element is a column of a complete table in the context information, but the second display content only contains a portion of the columns of that table, the second display content can be expanded to add the missing column; if the target image element is the complete outline of an object in the context information, but the second display content includes the object and surrounding clutter, the clutter can be deleted, leaving only the object.
[0091] In one example, refer to Figure 5 If a user selects the table comparing model performance in the reading document, but the model structure diagram in the context information shows the structure of the model, and the correlation between model structure and model performance is 0.9, which is greater than the target threshold of 0.8, it indicates that the model structure diagram and model performance comparison table are of interest to the user. Therefore, the model structure diagram and model performance comparison table will be used as the target analysis results.
[0092] In another example, refer to Figure 4 If a user selects the kitten, but the context describes the kitten playing with a ball, the correlation between the ball and the kitten is calculated to be 0.85, which is greater than the target threshold of 0.8. This indicates that both the ball and the kitten are of interest to the user, and therefore, the ball and the kitten are considered as the target analysis results.
[0093] In other embodiments, if the context information includes text content and image content, the semantic description corresponding to at least one image element in the context information and the target image element are identified; based on the semantic description and the target image element, at least one image element in the image content corresponding to the second display content is edited to obtain the target analysis result.
[0094] In one example, a user wants to select the main product image on an e-commerce page displayed on the screen. The page displays a large main product image (a blue backpack) at the top, followed by four smaller detailed images below (close-up of the backpack's zipper, internal structure, side view, and shoulder strap details). The right side of the page includes the text description, "This backpack is made of waterproof fabric and features thickened shoulder straps and multi-functional compartments." The user uses a touchscreen to freely select on this page using a curved path. Due to the large area of the selection, the first area formed by the first trajectory not only covers the main product image but also includes the zipper close-up and internal structure images from the four detailed images below, as well as part of the text description "waterproof fabric." The second display content corresponding to the first area (i.e., the main product image, zipper close-up, internal structure image, and "waterproof fabric" text) and the first display content (i.e., the complete content of the e-commerce page) are input into the target processing model.
[0095] The target processing model first performs layout analysis on the first displayed content, identifying the main product image area at the top of the page, the four detail images below, and the text description area on the right. The target processing model extracts the text description "This backpack is made of waterproof fabric and features thickened shoulder straps and multi-functional compartments" as the text content in the context information, and simultaneously extracts the four detail images below as the image content in the context information.
[0096] Based on the text content within the context information, the target processing model identifies the semantic descriptions corresponding to each image element in the second displayed content. The semantic description corresponding to the main product image is "this backpack," the semantic description corresponding to the zipper close-up image is "multi-functional compartment" (implicit association), and the semantic description corresponding to the internal structure diagram is "multi-functional compartment." The target processing model analyzes the semantic descriptions and determines that the text "waterproof fabric" is a description of the product material rather than a direct description of the image element. Therefore, this text does not belong to the image content the user expects to select, and based on this semantic description, the text "waterproof fabric" is removed from the second displayed content.
[0097] Based on the image content in the context information, the target processing model identifies the four detail images below as target image elements. The model calculates the similarity and relevance between these target image elements and each image element in the second display content. The main product image and the four detail images have high color similarity (all are blue backpacks), but significant size differences. The zipper close-up image's positional relevance to the main product image is shown as "the detail image is below the main image," and its action relevance is shown as "the zipper is part of the backpack." The internal structure image also has positional and action relevance to the main product image. The target processing model sets the target threshold to a similarity of 0.6 and a relevance score of 0.5 or higher. The calculation results show that although the zipper close-up image and the internal structure image have some relevance to the main product image, as independent detail display images, they belong to different display units semantically from the user-selected main product image, with a relevance score of 0.4, below the target threshold. Based on this judgment, the target processing model edits the image elements in the second display content: deleting the zipper close-up image and the internal structure image, retaining only the main product image, and generating a target analysis result containing only the main product image. The area containing the main product image is designated as the second area, serving as the user's final selection result and allowing for subsequent operations such as saving to the local album, sharing with friends, and comparing product prices.
[0098] Understandably, by combining the text and image content in the context information to perform multi-dimensional analysis and editing of the image elements in the second display content, it is possible to accurately distinguish between the core image content that the user truly intends to select and the auxiliary image or text content that has been mistakenly selected.
[0099] As described above, in operation S240, a second region is determined based on the target analysis results. In one possible implementation, this operation may further include the operation of determining the second region based on the target analysis results if the confidence level corresponding to the target analysis results is greater than or equal to a preset confidence threshold.
[0100] For example, confidence level can be a quantitative assessment of how reliable the target processing model is of its output target analysis results. Confidence level is usually expressed in numerical form, such as a percentage or a probability value between 0 and 1, with higher values indicating greater confidence in the output result.
[0101] A preset reliability threshold can be a critical value used to determine whether to adopt the target analysis results. For example, the preset reliability threshold could be 0.8, 0.85, or 0.9. The preset reliability threshold can also be dynamically adjusted according to the application scenario; for example, it could be set to 0.9 in a high-precision document editing scenario and 0.7 in a fast browsing scenario.
[0102] In one example, a user wants to select the third paragraph on page seven of a technical report displayed on the screen. The user drags a selection box around the paragraph with the mouse. Due to the offset of the mouse starting point, the first region not only covers the target paragraph but also includes the blank area above it and the first line of text of the paragraph below, "Chapter Summary: Based on the above analysis,". The report page seven with the user's selection trajectory is input into the target processing model. The model identifies the main content within the user's selected area as the complete third paragraph and outputs, "Third paragraph: Experimental results show that this method has significant advantages in handling complex scenarios… In summary, this method has good practicality." The output confidence level is 0.91. Since 0.91 is greater than the preset confidence threshold of 0.8, the system judges that the target processing model's understanding of the user's intent is highly reliable and directly determines the second region based on the target analysis results. The system determines the precise boundary area corresponding to the third paragraph as the second region, automatically excluding the blank area above and the first line of text of the paragraph below that the user mistakenly selected, as the user's final selection result, which can be used for subsequent operations such as copying, highlighting, and adding annotations.
[0103] In another possible implementation, the operation may further include the operation of determining a second region based on the second displayed content if the confidence level is less than a preset confidence threshold.
[0104] In one example, a user wants to select a specific icon from a design drawing displayed on the interface. The user uses a stylus to make a free-form selection on the drawing, covering most of the target icon's area. However, due to the complex background and the presence of multiple similar icons, the first area includes not only the target icon but also parts of the background grid lines and the edges of adjacent icons. When the design drawing with the user's selection trajectory is input into the target processing model, due to the high icon similarity and severe background interference, the target analysis model outputs a "combination area containing three icons," with a confidence level of only 0.42. Since 0.42 is less than the preset minimum confidence threshold of 0.5, the system determines that the target processing model's understanding of the user's intent is significantly flawed and discards the target analysis result. The system detects that the user's original selection trajectory is a non-closed curve, with the starting point coordinates (230, 410) and the ending point coordinates (235, 408), and the distance between the two points is less than 10 pixels. The system forms a closed trajectory by connecting the starting and ending points, and then smooths the trajectory, calculating the minimum bounding rectangle of the smoothed trajectory with boundary coordinates (220, 400, 340, 480). The system designates this rectangular area as the second region, preserving the user's original selection intent, and uses it as the user's final selection result.
[0105] In another possible implementation, the operation may further include the operation of determining a second region based on the target analysis results and the second displayed content if the confidence level is within a preset confidence level range.
[0106] For example, the preset reliability range can be a numerical interval between a preset reliability threshold and a preset minimum confidence threshold. For instance, if the preset reliability threshold is 0.8 and the preset minimum confidence threshold is 0.5, then the preset reliability range can be 0.5 to 0.8. The preset reliability range can also be a combination of multiple discrete intervals, such as 0.4 to 0.6 and 0.7 to 0.8.
[0107] In one example, the user wants to select the fifth page of the academic paper displayed on the screen. Figure 2 And its corresponding captions. Users utilize the touchscreen to... Figure 2 The selection path is irregular due to the instability of the finger swipe, and the first area is covered. Figure 2 Most of the area and the first half of the caption "Fig. 2. The architecture of". The fifth page of the paper with the user-selected trajectory is input into the target processing model, and the target processing model outputs " Figure 2 The system outputs a confidence level of 0.68. Since 0.68 falls within the preset confidence level range of 0.5 to 0.8, the system determines that it needs to comprehensively analyze the target analysis results and the user's original selection of the second display content. Based on the target analysis results and the second display content, the system will... Figure 2 ,and Figure 2 The area corresponding to the caption text is used as the second area, which serves as the user's final selection result.
[0108] Understandably, by introducing a confidence judgment mechanism, the system can dynamically select a strategy that relies entirely on model inference, completely retains the user's original choice, or combines the advantages of both, based on the reliability of the target processing model's understanding of the user's intent. This ensures selection accuracy while taking into account user operating habits, avoiding selection bias caused by model misjudgment and preventing content loss or redundancy caused by inaccurate user operation. It achieves a balance between intelligent correction and user control, further enhancing the robustness and adaptability of the selection function.
[0109] As described above, the operation involves determining a second region based on the second displayed content. In one possible implementation, this operation may further include the operation of: if the first trajectory is a first type of trajectory, editing the first trajectory based on its morphological parameters to determine the second region.
[0110] If the first trajectory is a second type of trajectory, determine the semantic integrity of the second displayed content, edit the first trajectory based on the semantic integrity, and determine the second region.
[0111] The first type of trajectory differs from the second type of trajectory.
[0112] For example, a first-type trajectory can be a trajectory path where the starting and ending points fail to connect to form a closed shape when the user performs a selection operation. For instance, if a user draws a circle around a table with a stylus, and the starting point coordinates are (100, 200) and the ending point coordinates are (105, 198), with a distance of 5.4 pixels between the two points, and no closure is formed, then this trajectory is a non-closed trajectory, i.e., a first-type trajectory.
[0113] Editing can be a process of correcting, completing, or optimizing the second displayed content or the first trajectory. Editing includes, but is not limited to, trajectory closure, boundary expansion, and content semantic alignment. The purpose of editing is to transform imprecise user selections into accurate region boundaries, ensuring that the second region fully covers the content the user intends to select.
[0114] The second type of trajectory can be a trajectory used to annotate or emphasize content. The second type of trajectory can include, but is not limited to: rectangles, underlines, wavy lines, straight lines, arrows, etc.
[0115] In one example, a user wants to select Table 3 on page 10 of the product manual displayed on the screen. The user uses a stylus to select around Table 3. Due to the stylus moving too quickly and the user's hand shaking, the starting point of the selection trajectory is (150, 350), and the ending point is (162, 345), with a distance of 13.9 pixels between the two points, exceeding the preset closing distance threshold of 10 pixels, thus forming a non-closed trajectory. The first trajectory covers the second display content, including the complete header row, data rows, and most of the table's outer border of Table 3. However, because the trajectory is not closed, there is a gap of approximately 15 pixels in the lower right corner. The system detects that the selection operation is a selection operation and that the first trajectory is non-closed, triggering the editing process. The system first calculates the boundary clarity of the first trajectory: by detecting the distance deviation between the trajectory and the outer border of Table 3, it finds that the average distance between the trajectory and the table border on the left, top, and bottom sides is 3 pixels, 4 pixels, and 5 pixels respectively, with relatively small deviations; however, the trajectory as a whole exhibits obvious jagged jitter, with a smoothness score of 0.62. The calculated boundary sharpness is 0.58, which is less than the preset sharpness threshold of 0.7. The system determines that the boundary of the user-drawn first trajectory is not clear enough and decides to adjust it based on the semantic boundary of the second displayed content. Through semantic analysis of the second displayed content, it identifies its main content as the complete table structure of Table 3, with the semantic boundary being the standard outer frame of the table, and boundary coordinates of (145, 345, 520, 580). The system defines the rectangular area corresponding to this semantic boundary as the second region, automatically correcting the non-closed gaps and jitter deviations caused by the user's selection, ensuring that the second region completely covers Table 3, serving as the user's final selection result.
[0116] In another example, a user wants to select the phrase "birds are flying in the sky" in a document. The user uses a stylus to underline the phrase, but due to the randomness of handwriting, the underline starts below "at" and ends below "fly." It was found that the underlined content "flying in the sky" was semantically incomplete. Based on semantic completeness, the underline was expanded to fully annotate "birds are flying in the sky," and the area containing "birds are flying in the sky" was designated as a second area.
[0117] As described above, the operation involves determining a second region based on the target analysis results and the second displayed content. In one possible implementation, this operation may further include the operations of determining a fusion strategy based on the target analysis results and the second displayed content; and determining the second region based on the fusion strategy. The fusion strategy includes at least one of the following: calculating the intersection of the regions corresponding to the target analysis results and the second displayed content; calculating the union of the regions corresponding to the target analysis results and the second displayed content; and determining the associated regions of the regions corresponding to the target analysis results and the second displayed content.
[0118] For example, the fusion strategy may be to combine the region corresponding to the target analysis result with the region corresponding to the second display content to determine the calculation rules for the final second region.
[0119] The intersection of the target analysis result and the area corresponding to the second displayed content is calculated to determine the overlapping part of the two areas as the second area. If the area corresponding to the target analysis result is "birds eating insects" and the area corresponding to the second displayed content is "eating insects to cure diseases", and the two areas have an overlapping part "eating insects", then by calculating the intersection, the second area is the area where "eating insects" is located.
[0120] The union of the regions corresponding to the target analysis result and the second displayed content is calculated, representing the merging of the two regions into a larger region as the second region. For example, if the region corresponding to the target analysis result is "birds flying in the sky" and the region corresponding to the second displayed content is "flying in the sky", the two regions partially overlap but each contains content not covered by the other. By calculating the union, the second region is "birds flying in the sky", which covers all the content of both regions.
[0121] Determining the associated region between the target analysis result and the second displayed content means inferring a new region as the second region based on the spatial or semantic relationship between the two regions. For example, if the region corresponding to the target analysis result is located in the upper 30% of the screen and the region corresponding to the second displayed content is located in the lower 30% of the screen, and the two regions do not have a direct inclusion relationship or overlap, by analyzing the spatial boundaries of the two regions, we can use the lower boundary of the target analysis result and the upper boundary of the second displayed content region to infer that the user intends to select the region located between these two boundaries, that is, the middle 30% of the screen. Thus, this middle region is determined as the associated region and serves as the second region.
[0122] Figure 6 The diagram illustrates a fourth schematic of a data processing method according to an embodiment of the present disclosure.
[0123] As described above, the data processing method of this embodiment may further include at least one of the following operations: if the selection operation includes a selection operation, based on the second region and the first trajectory, adjust the selection range corresponding to the first trajectory to obtain the second trajectory; highlight the selection range corresponding to the second trajectory on the display interface.
[0124] If the selected operation includes a sliding operation, based on the second region and the first trajectory, adjust the starting point and / or ending point corresponding to the first trajectory to obtain the second trajectory; highlight the second trajectory on the display interface.
[0125] Highlight the second area on the display interface by at least one of the following: adding a background color, highlighting a border, adding a semi-transparent mask, or zooming in on the second area.
[0126] In one example, refer to Figure 6 The user selected a portion of the content related to "the effect of drug A in treating hypertension" in the displayed document (the area enclosed by the first trajectory). The user's selection and the content of the displayed document were input into the target processing model to obtain contextual information about the user's selection. Among this contextual information, "side effects of drug A in treating hypertension" is essential information regarding the effect of drug A in treating hypertension. The correlation between "side effects of drug A in treating hypertension" and "the effect of drug A in treating hypertension" is 0.9, which is greater than the target threshold of 0.8. This indicates that "side effects of drug A in treating hypertension" and "the effect of drug A in treating hypertension" are of interest to the user. Therefore, the area containing the content corresponding to "side effects of drug A in treating hypertension" and "the effect of drug A in treating hypertension" is designated as the second region, and the first trajectory is adjusted to obtain the second trajectory. The content related to "side effects of drug A in treating hypertension" and "the effect of drug A in treating hypertension" is then highlighted in red on the display interface.
[0127] As described above, in operation S240, based on the target analysis results, a second region is determined. In one possible implementation, this operation may further include: determining the confidence level of the target analysis results; based on the confidence level, determining a first weight of the target analysis results and a second weight of the second displayed content; the confidence level is positively correlated with the first weight and negatively correlated with the second weight; and determining the second region based on the target analysis results, the first weight, the second displayed content, and the second weight.
[0128] For example, the first weight can be a proportionality coefficient assigned to the impact of the target analysis results in the process of determining the second region. The value of the first weight can be between 0 and 1. For instance, if the confidence level is 0.88, the corresponding first weight can be set to 0.85, indicating that the final determined second region will be mainly based on the semantic boundary inferred by the target processing model.
[0129] The second weight can be a proportional coefficient assigned to the influence of the second display content (i.e., the content originally selected by the user) during the determination of the second region. The value of the second weight can range from 0 to 1. The sum of the first weight and the second weight can be equal to 1 to ensure the completeness of the weight allocation. For example, if the first weight is 0.85, the corresponding second weight is 0.15, indicating that the finally determined second region will retain a small amount of the geometric features originally selected by the user.
[0130] The positive correlation between confidence level and first weight can be understood as follows: as the confidence level increases, the first weight also increases. For example, if the confidence level increases from 0.6 to 0.9, the corresponding first weight may increase from 0.55 to 0.88. This positive correlation reflects that the more confident the model is in its judgment, the more likely the system is to adopt the model's intelligent inference.
[0131] The negative correlation between confidence level and the second weight can be understood as follows: as the confidence level increases, the second weight decreases. For example, if the confidence level increases from 0.6 to 0.9, the corresponding second weight may decrease from 0.45 to 0.12. This negative correlation ensures that when the model's judgment is uncertain, the system will respect the user's original operational intent more. In one example, a user wants to select the second paragraph in a webpage article displayed on the screen. The user drags a selection box on the webpage with the mouse, but due to imprecise operation, the resulting rectangle includes most of the second paragraph, the last two lines of the first paragraph, and the first line of the third paragraph. A screenshot of the webpage with the user's selection trajectory is input into the target processing model. The target processing model identifies that the user's selection area mainly covers the second paragraph, infers that the user's intent is to select the complete second paragraph, and outputs a target analysis result containing the complete second paragraph, along with a confidence level of 0.76. Based on the confidence level of 0.76, the system calculates a first weight of 0.72 and a second weight of 0.28. The target analysis result corresponds to the boundary rectangle of the complete second paragraph, with coordinates (100, 250, 800, 420); the user's original selection corresponds to the first region with coordinates (120, 210, 780, 450). The system performs weighted fusion calculations: the x-coordinate of the top left corner of the second region = 0.72 × 100 + 0.28 × 120 = 105.6, the y-coordinate of the top left corner = 0.72 × 250 + 0.28 × 210 = 238.8, the x-coordinate of the bottom right corner = 0.72 × 800 + 0.28 × 780 = 794.4, and the y-coordinate of the bottom right corner = 0.72 × 420 + 0.28 × 450 = 428.4. The final coordinates of the second region are (106, 239, 794, 428). This region contains the complete main content of the second paragraph and also retains the geometric features of the user's original operation on the boundary. It serves as the user's final selection result and can be used for subsequent operations, such as copying text, generating summaries, and translating paragraphs.
[0132] Understandably, by introducing confidence level as a dynamic adjustment factor for weight allocation, the influence ratio between model inference and user's original operation is intelligently balanced. This allows the intelligent error correction capability to be fully utilized when the model has high confidence, and the user's operation intention to be effectively preserved when the model has low confidence, thus achieving a flexible and reliable region determination mechanism.
[0133] In some embodiments, determining the confidence level of the target analysis result includes at least one of the following: performing semantic analysis on the second display content; if the number of semantics indicated by the second display content is greater than or equal to the target number threshold, the second weight is less than the target weight threshold; if the number of semantics indicated by the target analysis result is greater than or equal to the target number threshold, the first weight is less than the target weight threshold.
[0134] For example, the number of semantic units can be the number of independent semantic units identified through semantic analysis. The number of semantic units reflects the semantic complexity or semantic diversity of the displayed content. For instance, if the second displayed content is "the user needs to repeatedly adjust the selection range," semantic analysis identifies two semantic units: the first is "adjust the selection range"; the second is "repeatedly adjust the selection range."
[0135] The target number threshold can be a preset value used to determine whether the semantic complexity is too high. The target number threshold can be set according to the application scenario, for example, it can be set to 2, 3 or other positive integers. For example, when the target number threshold is set to 3, it means that if the number of identified semantic units reaches or exceeds 3, the system will determine that the semantics of the content is too complex and there may be multiple unrelated semantic units mixed together.
[0136] The target weight threshold can be a preset upper limit used to limit the weight value. The target weight threshold can be a value between 0 and 1, such as 0.5, 0.6, or other values. For example, if the target weight threshold is set to 0.5, it means that when a weight is determined to need to be reduced, the value of that weight should be less than 0.5.
[0137] If the number of semantic units indicated by the second displayed content is greater than or equal to the target number threshold, and the second weight is less than the target weight threshold, this can be understood as follows: when the second displayed content selected by the user contains too many independent semantic units, the system reduces its reliance on the user's original selection. For example, if the user accidentally selects three unrelated paragraphs, resulting in the second displayed content containing 5 independent semantic units, and the target number threshold is set to 3, the system determines that the user's original selection has a significant "multiple selections" problem. In this case, the second weight is reduced to 0.3 (less than the target weight threshold of 0.5), and the first weight is correspondingly increased to 0.7, making the system rely more on the intelligent inference of the target processing model to exclude redundant content.
[0138] If the number of semantic units indicated by the target analysis result is greater than or equal to the target number threshold, and the first weight is less than the target weight threshold, this can be understood as follows: when the analysis result output by the target processing model itself contains too many independent semantic units, the system reduces its reliance on the model's inference results. For example, if the target processing model, when analyzing user selections, outputs target analysis results containing four semantic units from different topics due to complex contextual information, and the target number threshold is set to 3, the system determines that the model's inference may have over-expansion or unclear semantic boundaries. In this case, the first weight is reduced to 0.4 (less than the target weight threshold of 0.5), and the second weight is correspondingly increased to 0.6, allowing the system to retain more of the geometric range of the user's original selection and avoid excessive model intervention.
[0139] For example, the target analysis results and the user's original selections can be input into a new processing model to generate the final, accurate selection boundaries. Here, the new processing model can be a large language model or a multimodal model (VLM). Through pre-training, the model learns how to balance the first weight of the target analysis results ( The second weight of the user's original selection content () and the second weight of the content () ).
[0140] The second trajectory of the second region can then be represented as: ,in, It could be the trajectory of the displayed content in the second region predicted by the model. It could be the initial trajectory of the user's original selection of content. .
[0141] It is not a fixed value, but is dynamically generated by VLM based on the input features.
[0142] If the VLM has high confidence in the user's intent (e.g., the selected region clearly falls within a complete semantic unit, and the contextual information strongly supports a certain intent), then It will be higher, with VLM's predictions dominating.
[0143] If the VLM has low confidence in the user's intent (e.g., the selection region is very ambiguous, or there are multiple possible semantic interpretations), or the user's original choice is very close to a semantic boundary, then W USER It will be relatively higher, and tends to retain some characteristics of the user's original selection.
[0144] VLM can be trained to learn when to rely more on AI's understanding and when to respect the user's original input.
[0145] When the VLM outputs a new prediction boundary, the rendering engine immediately updates the highlighted areas on the screen. Double buffering is employed to ensure a smooth, flicker-free rendering process. The rendered result can be directly overlaid on the original screen content, for example, using semi-transparent fills or drawing precise borders, giving the user the visual experience of a "selection box intelligently snapping" to the target content. This requires highly optimized rendering code that fully utilizes hardware acceleration.
[0146] The precise selections determined by VLM are used as standard inputs and directly passed to various downstream intelligent applications (such as translation, summarization, OCR text copying to clipboard, image recognition, intelligent search, etc.), or exported as standard data formats.
[0147] Standardized API interfaces, such as JSON (JavaScript Object Notation) and Protobuf (Protocol Buffers), can be provided so that other applications or services can easily integrate and call this smart selection function.
[0148] In one example, a user wants to select the paragraph about "data encryption algorithms" in a technical document displayed on the screen. The user uses the touchscreen to make free-form selections on the document page, but due to the excessive range of the gestures, the selection path not only covers the target paragraph but also includes the "Network Protocol Overview" paragraph above and the "User Authentication Process" paragraph below. The document page with the user's selection path is input into the target processing model. The model performs semantic analysis on the second displayed content, identifying three independent technical topics: "Network Protocols," "Data Encryption," and "User Authentication," with a semantic count of 3. The system's target count threshold is 2. The model determines that the semantic count of the second displayed content (3) is greater than the target count threshold (2), indicating a significant multi-circle problem in the user's original selection. According to the rules, the system sets the second weight to 0.25 (less than the target weight threshold of 0.5). Simultaneously, based on the document's layout structure and semantic coherence analysis, the target processing model infers that the user's true intention is to select content related to "data encryption algorithms," outputting a target analysis result containing only the paragraph with this topic. This result has a semantic count of 1, less than the target count threshold, therefore the first weight is set to 0.75. The target analysis result corresponds to the boundary rectangle of the "Data Encryption Algorithm" paragraph, with coordinates (150, 320, 750, 480). The first region corresponding to the user's original selection has coordinates (140, 180, 760, 620), spanning three paragraphs. The system performs a weighted fusion calculation: the x-coordinate of the top left corner of the second region = 0.75 × 150 + 0.25 × 140 = 147.5, the y-coordinate of the top left corner = 0.75 × 320 + 0.25 × 180 = 285, the x-coordinate of the bottom right corner = 0.75 × 750 + 0.25 × 760 = 752.5, and the y-coordinate of the bottom right corner = 0.75 × 480 + 0.25 × 620 = 515. The final coordinates of the second region are (148, 285, 753, 515). This region mainly covers the "Data Encryption Algorithm" paragraph, effectively eliminating the other two topic paragraphs that the user mistakenly selected. At the same time, it slightly retains the range characteristics of the user's original operation on the boundary, serving as the user's final selection result and allowing for subsequent operations.
[0149] In some embodiments, determining the second region includes: if the target analysis result indicates two semantic boundaries, determining the boundary characters corresponding to the semantic boundaries; calculating the coverage ratio of the first trajectory on the display area of the boundary characters; if the coverage ratio is greater than or equal to a target coverage threshold, taking the region corresponding to the semantic boundary containing the boundary characters as the second region; if the coverage ratio is less than the target coverage threshold, taking the region corresponding to the semantic boundary not containing the boundary characters as the second region.
[0150] Exemplarily, the two semantic boundaries can be two possible semantic range boundaries identified by the target processing model when analyzing the second display content. The two semantic boundaries usually correspond to the "narrow boundary" and "wide boundary" of the content, or the "core boundary" and "extended boundary". For example, if the user selects the text area in the document that contains "BAC, the CEO of Company A, announced a new strategy at the annual general meeting of shareholders", the target processing model may identify two semantic boundaries: the first boundary only contains the core person name "BAC", and the second boundary contains the complete subject-predicate structure "BAC announced a new strategy at the annual general meeting of shareholders". The two semantic boundaries can also correspond to semantic units of different granularities. For example, the first boundary corresponds to a single sentence, and the second boundary corresponds to the complete paragraph containing the sentence.
[0151] The boundary character can be a specific literal character that identifies the starting or ending position of the semantic boundary. The boundary character is usually the first or last character of the semantic unit. For example, if the first semantic boundary is "BAC", the corresponding boundary characters can include the starting boundary character "B" and the ending boundary character "C"; if the second semantic boundary is "BAC announced a new strategy at the annual general meeting of shareholders", the corresponding boundary characters can include the starting boundary character "B" and the ending boundary character "略". The boundary character can also be a punctuation mark, such as a period, comma, etc., as a marker for semantic separation.
[0152] The display area of the boundary character can be the visible pixel area occupied by the character on the display interface. For text characters, the display area is usually the smallest rectangle enclosing the character. For example, if the boundary character "略" is presented in the 14th font on the display interface, its display area can be a rectangular area with coordinates (520, 340, 534, 354), a width of 14 pixels, and a height of 14 pixels.
[0153] The coverage ratio can be a quantitative indicator of the degree of overlap between the first trajectory and the display area of the boundary character. The coverage ratio can be obtained by calculating the percentage of the area of the display area of the boundary character covered by the first trajectory in the total area of the display area. For example, if the total area of the display area of the boundary character "略" is 196*196 pixels, and the user's first trajectory covers 150*150 pixels in this display area, the coverage ratio is 76.5%.
[0154] The target coverage threshold can be a preset proportional value used to determine whether the user intends to select a certain boundary character. The target coverage threshold can be a value between 0 and 1. For example, it can be set to 0.5, 0.6, 0.7 or other values. <00003]]
[0155] In one example, a user wants to select a quote from a news article displayed on the screen. The original text of the fifth paragraph of the news article is: "According to First News Agency, A Company CEO BAC stated at the Q1 2024 earnings conference that the company's cloud computing revenue increased by 35% year-on-year, and orders for artificial intelligence-related products reached a record high, with full-year revenue expected to exceed $60 billion." The user drags a selection box on the paragraph with the mouse, starting at the letter "A" in "A" and ending near the word "high" in "record high." However, due to a slight shift in position when the mouse is released, the actual ending point falls approximately 5 pixels to the right of the word "high." The news page with the user's selection trajectory is input into the target processing model. The model performs semantic analysis on the second displayed content, identifying that the selected area mainly covers a complex quotation and inferring two possible semantic boundaries: the first boundary is "Company A CEO BAC stated at the Q1 2024 earnings conference that the company's cloud computing revenue increased by 35% year-on-year, and orders for AI-related products reached a record high," which is semantically complete but relatively long; the second boundary is "Company A CEO BAC stated at the Q1 2024 earnings conference that the company's cloud computing revenue increased by 35% year-on-year," which is semantically relatively independent and shorter. The system determines the terminating boundary character of the first boundary to be "high," and the terminating boundary character of the second boundary to be "%." By analyzing the user's first trajectory, the system detects that the trajectory covers approximately 110*110 pixels in the display area of the character "high" (coordinates (685, 342, 699, 356), area 196*196 pixels), calculating a coverage ratio of 0.561. The system sets the target coverage threshold to 0.6. Since the coverage ratio of 0.561 is less than the target coverage threshold of 0.6, the system determines that the user did not intentionally include the content "historical high" in the selection range, but rather wanted to select content up to "35%". Therefore, the system uses the region corresponding to the second semantic boundary that does not contain the boundary character "high" as the second region, namely the text region corresponding to "Company A CEO BAC stated at the Q1 2024 earnings conference that the company's cloud computing business revenue increased by 35% year-on-year", with coordinates ranging from (125, 342, 672, 356). This second region avoids both "insufficient circles" caused by user operation errors (which might be truncated in the middle of the character "high" if the user's trajectory is followed exactly) and "excessive circles" caused by over-expansion (which would be caused by using the first boundary containing "historical high"). It serves as the user's final selection result and can be used for subsequent operations.
[0156] Based on the above data processing method, this disclosure also discloses an electronic device, which includes at least one processor and a target application running on the processor. The target application is capable of independently executing or calling at least one artificial intelligence model to perform the following operations: obtaining a first display content of a display interface; determining a first region corresponding to the selection operation in response to a user's selection operation on the first display content; obtaining the first region based on a first trajectory formed by the selection operation; analyzing and processing the second display content corresponding to the first region using a target processing model to generate a target analysis result; determining a second region based on the target analysis result; the second region including at least a portion of the first region.
[0157] For example, the target application can be a software program running on a processor, capable of providing intelligent content processing and interaction functions. The target application can be an intelligent agent or an artificial intelligence assistant. For instance, the target application could be an intelligent agent application called "Smart Reading Assistant." An artificial intelligence assistant can be a conversational assistant built on a large language model, capable of interacting with users through natural language and understanding various user commands and needs.
[0158] At least one artificial intelligence model may include an object processing model. This AI model can be a neural network model trained using deep learning, possessing capabilities such as image recognition, text understanding, and semantic analysis. For example, the object processing model can be a multimodal fusion model capable of simultaneously processing visual and textual features. Its input includes the coordinate sequence of the user-selected trajectory and the displayed content within a first region, while its output includes information such as semantic boundaries, confidence scores, and coordinates of the recommended region.
[0159] The operational details can be found in the above description of the data processing methods, and will not be repeated here.
[0160] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0161] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Without departing from the scope of this disclosure, various substitutions and modifications can be made by those skilled in the art, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A data processing method, comprising: Obtain the first displayed content of the display interface; In response to the user's selection operation on the first displayed content, a first area corresponding to the selection operation is determined; The first region is obtained based on the first trajectory formed by the selection operation; The target processing model is used to analyze and process the second display content corresponding to the first region to generate target analysis results; Based on the target analysis results, a second region is determined; the second region includes at least a portion of the first region.
2. The method according to claim 1, wherein the step of analyzing and processing the second display content corresponding to the first region using a target processing model to generate a target analysis result includes: The target processing model is used to analyze and process the first display content and the second display content to obtain the context information of the second display content; The second displayed content is edited using the context information to obtain the target analysis result.
3. The method according to claim 2, wherein the step of analyzing and processing the first display content and the second display content using the target processing model to obtain the context information of the second display content includes at least one of the following: Semantic analysis is performed on the first display content and the second display content, and the third display content in the first display content is used as the context information of the second display content; the semantic correlation between the third display content and the second display content is greater than a preset correlation threshold. A layout analysis is performed on the first display content and the second display content, and the fourth display content in the first display content is used as the context information of the second display content; the fourth display content and the second display content satisfy a preset positional relationship.
4. The method according to claim 2, wherein editing the second displayed content using the context information to obtain the target analysis result includes at least one of the following: The second display content is edited based on at least one fifth display content in the context information to obtain the target analysis result; the semantic correlation between the at least one fifth display content and at least a portion of the content in the second display content meets a preset condition; The second displayed content is edited based on the preset identifier in the context information to obtain the target analysis result.
5. The method according to claim 2, wherein editing the second displayed content using the context information to obtain the target analysis result includes: If the second displayed content includes image content, at least one image element in the image content is edited based on the context information to obtain the target analysis result; If the second displayed content includes image content, then editing at least one image element in the image content based on the context information to obtain the target analysis result includes at least one of the following: If the context information includes text content, identify the semantic description corresponding to the at least one image element in the context information; delete at least one image element that does not correspond to the semantic description, and obtain the target analysis result; If the context information includes image content, identify the target image element in the context information; Based on the target image element, at least one image element in the image content corresponding to the second display content is edited to obtain the target analysis result; the similarity and / or correlation between the target image element and at least one image element in the image content corresponding to the second display content is greater than the target threshold.
6. The method according to claim 1, wherein determining the second region based on the target analysis result includes: If the confidence level corresponding to the target analysis result is greater than or equal to a preset confidence threshold, the second region is determined based on the target analysis result; or If the confidence level is less than the preset confidence threshold, the second region is determined based on the second displayed content; or If the confidence level is within the preset confidence level range, the second region is determined based on the target analysis results and the second display content.
7. The method according to claim 6, wherein determining the second region based on the second displayed content comprises: If the first trajectory is a first type of trajectory, the first trajectory is edited based on its morphological parameters to determine the second region; If the first trajectory is a second type of trajectory, determine the semantic integrity of the second displayed content, and edit the first trajectory based on the semantic integrity to determine the second region; The first type of trajectory is different from the second type of trajectory.
8. The method according to claim 6, wherein determining the second region based on the target analysis result and the second display content includes: Based on the target analysis results and the second displayed content, a fusion strategy is determined; Based on the fusion strategy, the second region is determined; The fusion strategy includes at least one of the following: calculating the intersection of the target analysis result and the region corresponding to the second display content; calculating the union of the target analysis result and the region corresponding to the second display content; and determining the associated region of the target analysis result and the region corresponding to the second display content.
9. The method according to claim 1, further comprising at least one of the following: If the selection operation includes a selection operation, the selection range corresponding to the first trajectory is adjusted based on the second region and the first trajectory to obtain the second trajectory; the selection range corresponding to the second trajectory is highlighted on the display interface. If the selection operation includes a sliding operation, based on the second region and the first trajectory, the starting point and / or ending point corresponding to the first trajectory are adjusted to obtain the second trajectory; The second trajectory is highlighted on the display interface; Highlighting the second area on the display interface includes at least one of the following: adding a background color, a highlight border, a semi-transparent mask, or partial magnification to the second area.
10. An electronic device comprising at least one processor and a target application running on said processor, said target application being capable of independently executing or invoking at least one artificial intelligence model to perform the following operations: Obtain the first displayed content of the display interface; In response to a user's selection operation on the first displayed content, a first region corresponding to the selection operation is determined; the first region is obtained based on a first trajectory formed by the selection operation. The target processing model is used to analyze and process the second display content corresponding to the first region to generate target analysis results; Based on the target analysis results, a second region is determined; the second region includes at least a portion of the first region.