Text extraction method, electronic device, computer storage medium, and program product

CN122777030APending Publication Date: 2026-09-18DINGTALK (CHINA) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320469.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

目前的技术,仅支持针对单个文本组件内部的文本选取复制能力,但对于跨组件的文本复制却无能为力

Benefits of technology

[0008]According to a fourth aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the text extraction method as described in the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777030A_ABST
    Figure CN122777030A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a text extraction method, an electronic device, a computer storage medium and a program product, wherein the text extraction method comprises: in a display interface displaying dynamic information through a dynamic updating component, obtaining text in a plurality of components of a target region and text style information of the text; according to the text and the text style information, performing analog rendering on the text to obtain a rendering region corresponding to the text; according to a relationship between the text and the plurality of components, dividing the rendering region into a plurality of sub-regions, and sorting the plurality of sub-regions; and according to a sorting result, extracting text selected by a selection operation of a user from the plurality of components according to the selection operation. Through the scheme of the embodiments of the present application, cross-component text extraction can be effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a text extraction method, electronic device, computer storage medium, and computer program product. Background Technology

[0002] Currently, users frequently need to extract and copy text. For example, in chat and AI interaction scenarios, selecting and copying text are common and important functions. In such scenarios, text may exist in different interface components, such as text components, rich text components, image components, card components, and so on. Current technology only supports text selection and copying within a single text component, but it cannot handle text copying across components.

[0003] Therefore, how to conveniently extract and copy text across components, reduce the user's operational burden, and improve the user experience has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, embodiments of this application provide a text extraction scheme to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of the embodiments of this application, a text extraction method is provided, comprising: in a display interface that displays dynamic information through dynamically updated components, acquiring text in multiple components of a target area and text style information of the text; performing simulated rendering on the text based on the text and the text style information to obtain a rendering area corresponding to the text; dividing the rendering area into multiple sub-regions according to the relationship between the text and the multiple components, and sorting the multiple sub-regions; and extracting the text selected by the user's selection operation from the multiple components according to the sorting result.

[0006] According to a second aspect of the present application, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store a computer program; and the processor is used to execute the text extraction method described in the first aspect by running the computer program stored in the memory.

[0007] According to a third aspect of the embodiments of this application, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the text extraction method as described in the first aspect.

[0008] According to a fourth aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the text extraction method as described in the first aspect.

[0009] According to the text extraction scheme provided in this application embodiment, in a display interface that displays dynamic information through dynamically updated components, text and text style information of multiple components in a target area can be obtained. Then, based on the text and text style information, the text is simulated and rendered to obtain the corresponding rendering area. Afterwards, the rendering area is divided into multiple sub-regions according to the relationship between the text and multiple components, and these sub-regions are sorted. Based on the sorting result, the text selected by the user's text selection operation is extracted from multiple components. Therefore, on the one hand, the technical solution of this application embodiment can conveniently and effectively realize cross-component text selection and extraction from the display interface that displays dynamic information through dynamically updated components, effectively improving the user experience. On the other hand, since text style information is considered when determining the rendering area corresponding to the text in this application embodiment, the rendering area corresponding to the text can be obtained more accurately, facilitating more accurate cross-component text extraction from multiple components. Furthermore, since the sorting result obtained by sorting the multiple sub-regions of the rendering area according to the relationship between the text and multiple components can be used to extract the text selected by the user's text selection operation from multiple components, the selected text in the components can be extracted more accurately. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0011] Figure 1 This is a schematic diagram of a text extraction system according to an embodiment of this application.

[0012] Figure 2 This is a flowchart of the steps of a text extraction method according to an embodiment of this application.

[0013] Figure 3 This is a flowchart of an optional step in obtaining the rendering area corresponding to the text in an embodiment of this application.

[0014] Figure 4 This is an optional flowchart of step S2044 in an embodiment of this application.

[0015] Figure 5 This is a flowchart illustrating an optional step in sorting multiple sub-regions in an embodiment of this application.

[0016] Figure 6 This is a flowchart of an optional step in an embodiment of this application to extract the text selected by a selection operation from multiple components.

[0017] Figure 7 This is a schematic diagram illustrating an example of a text extraction scheme according to an embodiment of this application.

[0018] Figure 8 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0020] As mentioned earlier, in current scenarios such as chat and AI interaction, only the ability to select and copy text within a single text component is supported, but it is powerless to copy text across components.

[0021] Therefore, this application provides a text extraction scheme that enables free selection and copying of text across components, effectively improving the user experience.

[0022] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.

[0023] Figure 1 An exemplary system applicable to embodiments of this application is shown. For example... Figure 1 As shown, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106. Figure 1 The example in the text is 106 user devices.

[0024] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 can perform any suitable function, including but not limited to storing various application-related data, such as user interaction data from IM applications, etc.

[0025] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user equipment 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user equipment 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.

[0026] User device 106 may include any one or more user devices suitable for displaying information, interacting with users, etc. In some embodiments, user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include mobile devices, tablet computers, laptop computers, desktop computers, and / or any other suitable type of user device.

[0027] Optionally, when the user equipment 106 implements the scheme of the embodiments of this application, in some embodiments, the user equipment 106 can be used to perform a text extraction method. As an optional example, in some embodiments, the user equipment 106 can first obtain the text in multiple components of the target area and the text style information of the text in the display interface that displays dynamic information through dynamically updated components; then, it can simulate rendering the text according to the text and the text style information to obtain the rendering area corresponding to the text; then, according to the relationship between the text and multiple components, the rendering area is divided into multiple sub-regions, and the multiple sub-regions are sorted; then, according to the sorting result, the text selected by the user's selection operation is extracted from multiple components.

[0028] Based on the above system, this application provides a text extraction scheme, which will be described below through several embodiments.

[0029] Figure 2This is a flowchart illustrating the steps of a text extraction method according to an embodiment of this application. According to a first aspect of an embodiment of this application, a text extraction method is provided, referring to... Figure 2 As shown, the method includes steps S202, S204, S206, and S208, specifically:

[0030] S202: In a display interface that displays dynamic information through dynamically updated components, obtain the text and text style information of multiple components in the target area.

[0031] The aforementioned display interface can be any type of display interface, such as an application display interface, a webpage display interface, etc. The display interface can display multiple components, which can be arbitrary. Some components may include those displaying text-based information, such as text components (which can display text), rich text components (used to display rich text information including text), interactive card components (which can display interactive information including text in card form), image components (which can display images; some image components can also display text in addition to images), etc. Other components may not display text, such as video components (which can play videos), etc. By dynamically updating the components in the display interface to display dynamic information, the display interface can display rich content in real time. It should be understood that, for the technical solutions of this application embodiment, text selection and extraction can be performed for components that can display text.

[0032] The target area can be a specific region within the display interface. For example, a specific area can be defined from any position within the display interface using any shape (such as a selected rectangle). Alternatively, a specific sub-interface within the display interface can be used as the target area; this sub-interface could be a pop-up window or a fixed element within the display interface. In other embodiments, the target area can also be the entire display interface. In some possible implementations, the target area can be a default area, such as the currently displayed screen area; or a default area within the display area; or an area selected by the user; and so on.

[0033] Text style information indicates the style of text, and may include, but is not limited to, at least one of the following: text size, font, color, letter spacing, glyph, and font effects. Since different text styles may cause slight differences in the position and / or area occupied by the text, taking the text style information into account when determining the corresponding rendering area of ​​the text will yield more accurate information, thus facilitating more accurate cross-component text extraction from multiple components.

[0034] In practical applications, those skilled in the art can obtain text and text style information from multiple (two or more) components of a target area using any appropriate method. The specific implementation of step S202 is not limited in this embodiment. In some optional embodiments, step S202 may involve traversing the view tree of the target area of ​​the display interface, determining at least one view component within the target area, and obtaining text and text style information from multiple components of the target area from at least one view component. At least some of the multiple components are components of different types.

[0035] A view tree is a tree-like structure used to represent the structure of an interface and its rendered content. Through a view tree, one can obtain the various elements on the interface and their hierarchical relationships. In instant messaging (IM) applications, view trees can be used to efficiently manage and render user interface (UI) components. Each UI component can be considered a node in the view tree. These nodes are interconnected through parent-child relationships, forming a hierarchical structure. View trees can be implemented in any feasible way. Optionally, in native development, the view tree can be built using the operating system's UI framework (such as the View framework in Android or the UIView framework in iOS). Optionally, in cross-platform development, the view tree can be built using the rendering engine of the cross-platform development framework (such as React Native, Flutter, etc.). Optionally, if the IM application is developed using a web technology stack, the view tree can be implemented as a DOM (Document Object Model) tree, which can be constructed using HTML, CSS, and JavaScript. For example, a DOM tree can be constructed by parsing an HTML document, and CSS styles and JavaScript logic can be applied to render and update the interface. Each node represents an HTML element, and the relationship between nodes is defined through parent-child relationships. It should be understood that, in this embodiment, by obtaining the view tree of the target area of ​​the display interface, at least one view component of the target area can be accurately determined, and the text in multiple components within the target area, as well as the text style information of the text, can be accurately determined.

[0036] Alternatively, in order to accurately determine all text within the target area, all view components of the target area can be determined through the view tree.

[0037] Furthermore, in the embodiments of this application, the identified multiple components may be components of the same type, such as all being text components, but are not limited to this. It is also possible that at least some of the multiple components are components of different types. For example, different types of components may all include text, thereby meeting the need to extract text across components from multiple different types of components. For example, in some embodiments, the multiple components may include at least two types of components selected from text components, rich text components, interactive card components, and image components.

[0038] It should be understood that the embodiments of this application are not limited to using a view tree to obtain the text and text style information of multiple components in the target area, as long as the functional requirements can be met.

[0039] S204: Based on the text and text style information, simulate the rendering of the text to obtain the rendering area corresponding to the text.

[0040] In computer graphics, text simulation rendering can be the process of converting text information into images or visual effects through algorithms or models. In this embodiment, the simulation rendering is implemented in the background and is imperceptible to the user. Optionally, after obtaining the text in multiple components of the target area and the text style information, a preset text rendering engine can be used to simulate the rendering of the text and calculate the relevant information of the rendering area corresponding to the text. Any feasible text rendering engine can be used to implement this.

[0041] In some alternative embodiments, refer to Figure 3 The flowchart shown can be used to obtain the rendering area corresponding to the text through the following steps S2042 to S2044:

[0042] S2042: Based on the text and text style information, simulate the rendering of the text to determine the target rendering area of ​​the text within the target area, as well as the position information of the text within the target area.

[0043] Optionally, the text can be simulated and rendered using a preset text rendering engine based on the text and text style information. Based on the rendering result, the target rendering area of ​​the text within the target area and the text's position information within the target area can be determined. As an example, the text rendering engine can include, but is not limited to, the CoreText rendering engine and the DirectWrite rendering engine. However, those skilled in the art should understand that other rendering engines capable of simulating text rendering are also applicable to the solutions in this application's embodiments.

[0044] Optionally, the target rendering area of ​​the text within the target area can be calculated using the coordinates of the target rendering area in the coordinate system of the target area; the position information of the text within the target area can be calculated using the coordinates of the text in the coordinate system of the target area. Further optionally, the coordinates can be calculated using pixel coordinates.

[0045] S2044: Perform conversion processing based on the target rendering area and position information to obtain the rendering area of ​​the text in the display interface.

[0046] By transforming the target rendering area and the text's position information within the target area, the rendering area of ​​the text in the display interface can be accurately obtained, facilitating accurate cross-component text extraction from multiple components. Optionally, the rendering area of ​​the text in the display interface can be calculated using the coordinates of the text in the display interface's coordinate system, and these coordinates can be calculated in pixel coordinates.

[0047] In some alternative embodiments, refer to Figure 4 In the flowchart shown, step S2044 can be implemented as follows: S2044A to S2044B:

[0048] S2044A: Perform coordinate transformation on the target rendering area and position information to obtain the target rendering area and position information in the screen coordinate system.

[0049] In this application, coordinate transformation can be performed in any manner. For example, in some embodiments, a preset coordinate transformation formula can be used to calculate the target rendering area and position information, thereby achieving coordinate transformation to obtain the target rendering area and position information in the screen coordinate system.

[0050] In some alternative embodiments, the target rendering area and position information can be subjected to coordinate transformation based on the alignment and text configuration information between multiple components.

[0051] The alignment of components determines the arrangement of multiple components. For example, alignment can include, but is not limited to, at least one of the following: vertical center alignment, horizontal center alignment, left alignment, right alignment, justification, and distributed alignment.

[0052] Text configuration information can determine the state of the text within a component. For example, text configuration information may include, but is not limited to, configuration information such as line height, line spacing, and text indentation of the text within the component.

[0053] For example, when multiple components are aligned vertically to the center and the components have a line height set (e.g., 5mm), the relevant information of vertical center alignment and 5mm line height can be combined to perform coordinate transformation on the target rendering area and the position information of the text within the target area, so as to obtain the target rendering area and position information in the screen coordinate system.

[0054] It should be understood that since the target rendering area and position information obtained in step S2042 are usually preliminary areas and positions, the above optional embodiments in this application can accurately perform coordinate transformation processing on the target rendering area and position information according to the alignment of multiple components and the actual text configuration information, so as to accurately obtain the target rendering area and position information in the screen coordinate system, which is beneficial to subsequent accurate text positioning.

[0055] S2044B: Based on the target rendering area and position information in the screen coordinate system, obtain the rendering area of ​​the text in the display interface.

[0056] For example, the rendering area of ​​text in the display interface can be calculated based on the target rendering area and position information in the screen coordinate system.

[0057] Therefore, the technical solution of steps S2044A to S2044B in this embodiment can accurately and effectively determine the rendering area of ​​text in the display interface, which facilitates more accurate and convenient cross-component text extraction for multiple components.

[0058] S206: Divide the rendering area into multiple sub-regions based on the relationship between the text and multiple components, and sort the multiple sub-regions.

[0059] Optionally, the text rendering area in the display interface can be divided into multiple sub-regions based on the relationship between the text and multiple components, and sorted in any suitable way. For example, it can be sorted according to rules that conform to user habits. As an example, a "top-left-bottom-right" rule can be used, that is: sorting from the first sub-region located in the top left to the right, and then sorting each sub-region in turn until the last sub-region in the bottom right. This is more in line with users' usual reading habits and improves user acceptance. In addition, this sorting also facilitates cursor positioning during subsequent selection operations.

[0060] As an example for easier understanding, suppose there are a total of 6 sub-regions, with 2 sub-regions in each row. We can then sort them according to the "top-left-bottom-right" rule: the top-left sub-region (i.e., the leftmost sub-region in the first row) becomes the first sub-region, the rightmost sub-region in the first row becomes the second sub-region, the leftmost sub-region in the second row becomes the third sub-region, the rightmost sub-region in the second row becomes the fourth sub-region, the leftmost sub-region in the third row becomes the fifth sub-region, and the bottom-right sub-region (i.e., the rightmost sub-region in the third row) becomes the sixth sub-region. Of course, this is just an example; in practice, the arrangement can be more complex, but the general principle applies.

[0061] In other embodiments, sorting rules other than the "top left bottom right" rule can also be used. For example, sorting can be done by column, that is, sorting the first column first, then the second column, and so on until the last column is sorted. The settings can be configured as needed, and no restrictions are imposed here.

[0062] In some alternative embodiments, refer to Figure 5 The flowchart shown can be used to sort multiple sub-regions through the following steps S2062 to S2066:

[0063] S2062: On a component basis, determine that each of a plurality of components contains at least one line of target text.

[0064] In some embodiments of this application, segmentation rules can be set. Each component among multiple components can include one or more lines of text. It can be determined first that each component contains at least one line of target text.

[0065] Optionally, each line of text contained in each component can be identified as the target text.

[0066] S2064: Divide the rendering area into multiple sub-regions based on the part of each line of target text that corresponds to it in the rendering area.

[0067] Optionally, the corresponding part of each line of target text in the rendering area can be determined, and then these parts in the rendering area can be defined as sub-regions, instead of the other parts of the rendering area as sub-regions.

[0068] S2066: Sort multiple sub-regions based on their positional relationships.

[0069] Optionally, step S2066 can be understood by referring to the sorting rules such as "top left bottom right" and sorting by column introduced above, which will not be repeated here.

[0070] It should be understood that through the optional implementation of steps S2062 to S2066 above, the rendering area can be accurately and effectively divided into multiple sub-regions according to the relationship between the text and multiple components, and the multiple sub-regions can be accurately and effectively sorted according to the positional relationship between them, thereby facilitating subsequent text extraction.

[0071] S208: Based on the sorting results, extract the text selected by the user's text selection operation from multiple components.

[0072] Users can select text from multiple sorted sub-regions as needed. Based on the user's selection, the selected text can be extracted, realizing cross-component text extraction functionality for user convenience.

[0073] For example, the selection operation can be performed by at least one of the following: mouse selection, finger pressing / dragging selection, etc.

[0074] Based on this, the optional technical solutions of steps S204 to S208 in this embodiment of the application can, on the one hand, conveniently and effectively realize text selection and extraction of text from multiple components in the display interface of dynamically updated components displaying dynamic information, and allow cross-component text selection and extraction, effectively improving the user experience; on the other hand, since text style information is considered when determining the rendering area corresponding to the text in this embodiment of the application, the rendering area corresponding to the text can be obtained more accurately, which facilitates more accurate cross-component text extraction in the future; furthermore, since the sorting result obtained by sorting multiple sub-regions of the rendering area according to the relationship between the text and multiple components can be used to extract the text selected by the user's text selection operation from multiple components, the selected text in the components can be extracted more accurately and conveniently.

[0075] In some alternative embodiments, refer to Figure 6 The flowchart shown illustrates how steps S2082 to S2084 can be used to extract the text selected during the selection operation from multiple components.

[0076] S2082: Based on the sorting results, determine the target sub-region where the selected text is located from multiple sub-regions according to the user's selection operation on the text.

[0077] For example, when a user selects text by pressing and dragging with their finger, the user can long-press on a piece of text displayed on the touchscreen. The location of the long press can be obtained to determine the first target sub-region containing the selected text from multiple sub-regions. Then, based on the user's dragging operation and the sorting result of the multiple sub-regions, starting from the first target sub-region containing the selected text, the other target sub-regions containing the selected text can be determined sequentially until all target sub-regions containing the selected text have been determined.

[0078] For example, there are four sub-regions, numbered 1 through 4 (which can correspond to multiple components; for instance, sub-regions 1 and 2 correspond to component 1, and sub-regions 3 and 4 correspond to component 2). The text in sub-region 1 is "AXYZ", the text in sub-region 2 is "BCDE", the text in sub-region 3 is "FFFF", and the text in sub-region 4 is "GHIJ". If the user long-presses on the text "C" in the display, the second sub-region containing "C" is selected, and this second sub-region becomes the first target sub-region containing the selected text. If the user drags their finger to the text "H", then according to the sorting order, the third sub-region and the fourth sub-region containing "H" are also identified as target sub-regions containing the selected text. This is just a simple example; in reality, it may be more complex, but the analogy can be used to understand the concept.

[0079] S2084: Extract the text selected by the selection operation from the target component corresponding to the target sub-region among multiple components.

[0080] For example, using the above example of four sub-regions, if the identified target sub-regions are the 2nd, 3rd, and 4th sub-regions, then the selected text "CDE", "FFFF", and "GH" can be extracted from the target components corresponding to the 2nd, 3rd, and 4th sub-regions. This is just a simple example; in practice, it may be more complex, but the analogy can be used to understand the concept.

[0081] Based on this, in the optional embodiments of steps S2082 to S2084 of the present application, by determining the target sub-region where the selected text is located from multiple sub-regions according to the sorting result and the user's selection operation on the text, the text targeted by the selection operation can be accurately and conveniently extracted from the target components corresponding to the target sub-region among multiple components.

[0082] In some optional embodiments, step S208 can also be implemented as follows: based on the sorting results and the preset data structure for real-time recording of information in the sub-region, extract the text selected by the selection operation from multiple components according to the user's selection operation on the text.

[0083] It should be understood that this application embodiment designs a data structure for real-time recording of information in sub-regions, enabling real-time overall management of information in each sub-region. By using the sorting results and data structure, and based on the user's text selection operation, the selected text can be quickly extracted from multiple components, thereby improving text extraction efficiency.

[0084] The data structure described above can be implemented using any feasible data structure. For example, it can be implemented as an array, or it can be a custom data structure.

[0085] In this embodiment, the content recorded in the data structure can be set as needed. For example, in some optional embodiments, the information of the sub-region recorded in real time in the data structure includes at least: information about the region portion of the text in the sub-region corresponding to the region portion in the rendering region; information about the region portion of the selected text in the sub-region corresponding to the region portion in the rendering region; position information of the text in the sub-region; and content information of the text in the sub-region.

[0086] Optionally, the information about the region corresponding to the text within the sub-region in the rendering area may include: the positional information of the region corresponding to the text within the sub-region in the rendering area. This can be calculated using its coordinates in the coordinate system of the display interface.

[0087] Optionally, the information about the region corresponding to the selected text within the sub-region in the rendering area may include: the position information of the region corresponding to the selected text within the sub-region in the rendering area. This can be calculated using its coordinates in the coordinate system of the display interface.

[0088] Optionally, the position information of the text within the sub-region may include the specific position of each character in the text within the sub-region. This can be calculated using the coordinates of the text within the sub-region in the coordinate system of the display interface.

[0089] Therefore, the data structure provided in this application embodiment can facilitate the extraction of the text selected by the selection operation from multiple components by recording the above optional content, thereby improving the text extraction efficiency.

[0090] Optionally, the selected text within the sub-region can be filled with a background in the corresponding area of ​​the rendering region to visually display the selected text to the user, facilitating the user's selection operation. For example, the background fill can use any pattern or color; for instance, a solid color fill (or coloring) can be used. As an example, gray can be used for the background fill.

[0091] In some optional embodiments, the information of the selected text in the sub-region within the real-time recorded data structure corresponding to the region in the rendering area is obtained by: determining the starting cursor position and the ending cursor position of the selection operation based on the cursor drag position when the user selects the text; and obtaining the information of the selected text in the sub-region corresponding to the region in the rendering area based on the cursor drag position, the starting cursor position, and the ending cursor position.

[0092] Therefore, in this embodiment of the application, based on the cursor drag position, the starting cursor position, and the ending cursor position when the user selects text, the information of the area portion of the selected text in the sub-region corresponding to the rendering area can be accurately and effectively determined, so as to record it in the data structure in real time, so as to extract the selected text from multiple components in the future and improve the text extraction efficiency.

[0093] As explained earlier, selection can be performed through at least one of the following methods: mouse selection, finger pressing / dragging selection, etc. Therefore, determining the start cursor, the end cursor, and dragging the cursor can be controlled by the user using the mouse, by the user pressing and dragging the finger, or by any other feasible method.

[0094] For example, taking cursor control via user finger pressing and dragging as an example, suppose there are four sub-regions, numbered 1 to 4 (which can correspond to multiple components; for example, sub-regions 1 and 2 correspond to component 1, and sub-regions 3 and 4 correspond to component 2). The text in sub-region 1 is "AXYZ", the text in sub-region 2 is "BCDE", the text in sub-region 3 is "FFFF", and the text in sub-region 4 is "GHIJ". When the user drags the cursor, assuming the user first presses and holds the text "C" on the display interface, the starting cursor position can be determined to the left of text "C". If the user drags the cursor to text "H" and stops, the ending cursor position can be determined to the right of text "H". The cursor drag position can then be determined from text "C" to text "H". Afterwards, based on the cursor drag position, the starting cursor position, and the ending cursor position, the information of the selected text (i.e., texts "CDE", "FFFF", and "GH") within the sub-region can be obtained in the rendering area.

[0095] For example, alternatively, assuming the user's long-press position is not on the text (i.e., the long-press position cannot detect text), it can be considered a full selection operation. The starting cursor position can be determined as the left side of the first text in the first sub-region, and the ending cursor position can be determined as the last text in the last sub-region. The cursor drag position can be determined as the text in all sub-regions. For example, taking the above four sub-regions as an example, the starting cursor position can be determined as the left side of the text "A", the ending cursor position can be determined as the right side of "J", and the cursor drag position can be determined as from the text "A" to the text "J". Then, based on the cursor drag position, the starting cursor position, and the ending cursor position, the information of the area portion corresponding to the selected text (i.e., the selected text "AXYZ", "BCDE", "FFFF", "GHIJ") in the rendering area can be obtained.

[0096] It's understandable that for text within multiple sub-regions, unselected text can be either text before the starting cursor position or text after the ending cursor position. That is, during text selection, text can be categorized into three states: before the starting cursor region (e.g., text "AXYZ" and "B" in the previous example), after the ending cursor region (e.g., text "IJ" in the previous example), and selected (e.g., text "CDE", "FFFF", and "GH" in the previous example). The selected state can also be categorized into: starting cursor selected (e.g., text "C" in the previous example), fully selected (e.g., text "DE", "FFFF", and "G" in the previous example), and ending cursor selected (e.g., text "H" in the previous example). (The above can also be combined with...) Figure 7 (This can be understood using a scene diagram). Information about the relevant text can be recorded in real time using the data structure corresponding to the sub-regions, facilitating subsequent text extraction.

[0097] Optionally, the selected text within the sub-region calculated based on the starting cursor position, the cursor drag position, and the ending cursor position can be filled with background to visually display the selected text to the user, facilitating the user's selection operation. For example, the background fill can use any pattern or color; for instance, a solid color fill (or coloring) can be used. As an example, gray can be used for the background fill.

[0098] In some optional embodiments, when extracting the text selected by the selection operation from multiple components, the text selected by the selection operation can be extracted based on the information recorded in real time in the data structure corresponding to the sub-region where the selected text is located in the component selected by the selection operation.

[0099] It should be understood that since the selected text can be recorded in real time in the data structure corresponding to the sub-region, the information recorded in real time in the data structure corresponding to the sub-region where the selected text is located in the component selected by the selection operation in this embodiment can quickly and accurately extract the text selected by the selection operation, effectively improving the user experience.

[0100] For example, there are four sub-regions, numbered 1 to 4 (which can correspond to multiple components; for example, sub-regions 1 and 2 correspond to component 1, and sub-regions 3 and 4 correspond to component 2). The text in sub-region 1 is "AXYZ", the text in sub-region 2 is "BCDE", the text in sub-region 3 is "FFFF", and the text in sub-region 4 is "GHIJ". If a user long-presses on the text "C" in the display, the second sub-region containing "C" is selected, thus confirming that the second sub-region is the first sub-region containing the selected text. If a user drags their finger to the text "H" in a target sub-region, the selected text recorded in real-time by the data structure corresponding to the first sub-region will be empty (i.e., no selected text). The selected text recorded in real-time by the data structure corresponding to the second sub-region will be "CDE", the selected text recorded in real-time by the data structure corresponding to the third sub-region will be "FFFF", and the selected text recorded in real-time by the data structure corresponding to the fourth sub-region will be "GH". Therefore, based on the information recorded in real-time by the data structures corresponding to the four sub-regions, the selected text "CDE", "FFFF", and "GH" can be extracted. Of course, this is just a simple example; in reality, it may be more complex, but the analogy can be used to understand it.

[0101] In some optional embodiments, the information of the sub-region recorded in real time in the data structure may include at least one of the following: the head position of the sub-region; the tail position of the sub-region; the length information of the text in the sub-region; the width information of the text in the sub-region; and the height information of the text in the sub-region.

[0102] Optionally, the header position of the sub-region can be the position of the first character of the text within the sub-region. This header position can be calculated using the coordinates of the header position on the display interface.

[0103] Optionally, the end position of the sub-region can be the position of the last character of the text within the sub-region. This end position can be calculated using the coordinates of the end position on the display interface.

[0104] Optionally, the length information of the text within the sub-region can be the character length of the text within the sub-region.

[0105] Optionally, the width information of the text within the sub-region can be the character width of the text within the sub-region.

[0106] Optionally, the height information of the text within the sub-region can be the character height of the text within the sub-region.

[0107] Therefore, the data structure provided in this application embodiment records this optional information, which can facilitate the management of text in sub-regions and make it easier to accurately locate the text in sub-regions, so as to facilitate the extraction of text across components.

[0108] In some optional embodiments, a copy menu item function can also be set. After the user selects text, the text selected by the user can be extracted according to the text extraction method described above, and the extracted text can be copied to the clipboard so that the user can copy and paste the extracted text for use.

[0109] Optionally, one example application scenario of the text extraction scheme in this application embodiment is text extraction in IM applications on mobile terminals (e.g., including but not limited to mobile phones, tablets, etc.). For example, users can select text by pressing and dragging text on the display interface of the IM application, thereby achieving text extraction through the text extraction scheme in this application embodiment. For example, refer to the following... Figure 7 Understanding the process of creating a scene diagram.

[0110] For example, refer to Figure 7 The illustrated scenario provides an exemplary description of the implementation process of the text extraction scheme in this application embodiment. Figure 7As shown, in the display interface that displays dynamic information through dynamically updated components (taking the display interface of a mobile terminal's IM application as an example), there are two components (the first is denoted as component a, and the second as component b). Component a includes two texts, "AXYZ" and "BCDE", and component b includes two texts, "FFFF" and "GHIJ". Furthermore, the text style information (e.g., font size, font shape, character spacing, etc.) differs to some extent. Using the text extraction method of this embodiment, the text in multiple components of the target area and the text style information of the text can be obtained first (i.e., S202). Then, based on the text and text style information, the text is simulated and rendered to obtain the rendering area corresponding to the text (i.e., S204). Then, based on the relationship between the text and multiple components, the rendering area is divided into multiple sub-regions, and the multiple sub-regions are sorted (i.e., S206). Figure 7 As shown, four sub-regions p1, p2, p3, and p4 can be obtained. Based on the sorting results, and according to the user's text selection operations (e.g., long-pressing a text and then dragging to select other text), the text selected by the selection operation is extracted from multiple components (i.e., step S208). For example, the starting and ending cursor positions of the selection operation can be determined based on the cursor drag position when the user selects the text. Then, based on the cursor drag position, the starting cursor position, and the ending cursor position, the text between the starting and ending cursors is filled with background (which can be colored). The text selected by the selection operation, "CDE", "FFFF", and "GH", can be extracted based on the information recorded in real-time in the data structure corresponding to the sub-region where the selected text is located in each component selected by the selection operation. A copy function can then be provided to facilitate copying the extracted text. This enables cross-component text extraction and copying to meet user needs. It should be understood that... Figure 7 Further details of the implementation process can be understood in conjunction with the preceding embodiments, and Figure 7 The implementation process shown is merely an example for easy understanding of the embodiments of this application and is not intended to limit the embodiments of this application in any way.

[0111] In summary, the optional text extraction schemes in this application embodiment can, on the one hand, conveniently and effectively realize cross-component text selection and extraction from the display interface of dynamically updated components displaying dynamic information, effectively improving the user experience; on the other hand, since the text style information is considered when determining the rendering area corresponding to the text in this application embodiment, the rendering area corresponding to the text can be obtained more accurately, which facilitates more accurate cross-component text extraction in the future; furthermore, since the sorting result obtained by sorting the multiple sub-regions of the rendering area according to the relationship between the text and multiple components can be used to extract the text selected by the user's text selection operation from multiple components, the selected text in the components can be extracted more accurately and conveniently.

[0112] It is understandable that the technical solutions in related technologies can only support text selection of a single component. In complex pages and areas, only the text of a single component can be selected repeatedly, resulting in a poor user experience. However, as can be seen from the above description of the embodiments of this application, compared to related technologies, some technical solutions in the embodiments of this application adopt a self-implemented text region selection algorithm. By collecting text and planning regions, and dynamically calculating the selected region and text, it effectively solves the problem of poor user experience caused by the inability to select text in various complex pages and areas, or the limitation to selecting text from only a single component. In various display interfaces that dynamically update components to display dynamic information, in any complex area and scenario, the text extraction scheme in the embodiments of this application can be used to uniformly select, extract, and copy text, including but not limited to the selection, extraction, and copying of text in rich text components and interactive card components frequently used by users. It can achieve cross-component text selection and extraction for text from multiple components, which can significantly improve the user experience. The technical solution of this application embodiment can collect text, accurately calculate the text region range, design a text selection model, manage all text and selected region information in a unified manner, customize UI interaction, calculate and draw the selected region, realize cross-component text selection, extraction and copying, and provide a new approach to text extraction.

[0113] It is understood that the foregoing description of the text extraction method is merely an exemplary description of the embodiments of this application and is not intended to limit the embodiments of this application in any way.

[0114] According to a second aspect of the present application, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store a computer program; and the processor is used to execute the text extraction method described in the first aspect by running the computer program stored in the memory.

[0115] Figure 8 A structural block diagram of an optional electronic device according to an embodiment of this application is shown. This application does not limit the specific implementation of the electronic device 1000; however, as an example, reference is made to... Figure 8 The electronic device 1000 provided in this application embodiment includes: a processor 1002, a communication interface 1004, a memory 1006, and a communication bus 1008. Wherein:

[0116] The processor 1002, communication interface 1004, and memory 1006 communicate with each other via communication bus 1008.

[0117] Communication interface 1004 is used to communicate with other electronic devices or servers.

[0118] The processor 1002 is used to execute the computer program 1010, specifically the relevant steps in any of the aforementioned text extraction method embodiments.

[0119] Specifically, computer program 1010 may include program code that includes computer operation instructions.

[0120] The processor 1002 may be a CPU, a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0121] Memory 1006 is used to store computer program 1010. Memory 1006 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0122] Specifically, computer program 1010 can be used to cause processor 1002 to execute the text extraction method in any of the foregoing embodiments.

[0123] The specific implementation of each step in computer program 1010 can be found in the corresponding steps and units described in any of the foregoing text extraction method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0124] The electronic device 1000 in this application embodiment has been described in detail in the foregoing text extraction method embodiment. Therefore, its related content and beneficial effects can be understood by referring to the above method embodiment, and will not be repeated here.

[0125] According to a third aspect of the embodiments of this application, the embodiments of this application also provide a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the text extraction method described in any of the foregoing method embodiments. The computer storage medium includes, but is not limited to, compact disc read-only memory (CD-ROM), random access memory (RAM), floppy disk, hard disk, or magneto-optical disk, etc.

[0126] According to a fourth aspect of the embodiments of this application, the embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the text extraction method as described in any of the embodiments of the plurality of method embodiments described above.

[0127] The electronic device 1000 / computer storage medium / computer program product embodiment in this application has been described in detail in the foregoing text extraction method embodiment. Therefore, its related content and beneficial effects can be understood by referring to the above method embodiment, and will not be repeated here.

[0128] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0129] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0130] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0131] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for specific applications, but such implementations should not be considered beyond the scope of the embodiments of this application.

[0132] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". It should be noted that the concepts of "first", "second", etc., mentioned in the embodiments of this application are only used to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications of "a" and "a plurality" mentioned in the embodiments of this application are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0133] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. A text extraction method, comprising: In a display interface that displays dynamic information through dynamically updated components, the text in multiple components of the target area, as well as the text style information of the text, are obtained; Based on the text and the text style information, the text is simulated and rendered to obtain the rendering area corresponding to the text; Based on the relationship between the text and the multiple components, the rendering area is divided into multiple sub-regions, and the multiple sub-regions are sorted. Based on the sorting results, the text selected by the user's selection operation is extracted from the multiple components according to the selection operation of the text.

2. The method according to claim 1, wherein, The step of simulating rendering the text based on the text and the text style information to obtain the rendering area corresponding to the text includes: Based on the text and the text style information, simulated rendering of the text is performed to determine the target rendering area of ​​the text within the target area and the position information of the text within the target area; The text is rendered in the display interface by performing a conversion process based on the target rendering area and the location information.

3. The method according to claim 2, wherein, The step of performing conversion processing based on the target rendering area and the position information to obtain the rendering area of ​​the text in the display interface includes: The target rendering area and the position information are subjected to coordinate transformation to obtain the target rendering area and the position information in the screen coordinate system. The rendering area of ​​the text in the display interface is obtained based on the target rendering area in the screen coordinate system and the position information.

4. The method according to claim 3, wherein, The coordinate transformation process for the target rendering area and the position information includes: Based on the alignment method and text configuration information among the multiple components, coordinate transformation is performed on the target rendering area and the position information.

5. The method according to claim 1, wherein, The step of dividing the rendering area into multiple sub-regions based on the relationship between the text and the multiple components, and sorting the multiple sub-regions, includes: On a component-by-component basis, determine that each of the plurality of components contains at least one line of target text; The rendering area is divided into multiple sub-regions based on the portion of each line of target text that corresponds to it in the rendering area; The multiple sub-regions are sorted based on their positional relationships.

6. The method according to claim 5, wherein, The step of extracting the text selected by the user's selection operation from the multiple components according to the sorting result includes: Based on the sorting results, and according to the user's selection operation on the text, the target sub-region where the selected text is located is determined from the multiple sub-regions; The text selected by the selection operation is extracted from the target component corresponding to the target sub-region among the multiple components.

7. The method according to any one of claims 1-6, wherein, The step of extracting the text selected by the user's selection operation from the multiple components according to the sorting result includes: Based on the sorting results and the preset data structure for real-time recording of information in sub-regions, the text selected by the user's selection operation is extracted from the multiple components according to the user's selection operation on the text.

8. The method according to claim 7, wherein, The information recorded in real time for the sub-region includes at least: The information of the corresponding region portion of the text within the sub-region in the rendering area; The selected text within the sub-region corresponds to the area portion in the rendering area; The location information of the text within the sub-region; The text content information within the sub-region.

9. The method according to claim 8, wherein, The information recorded in real time for the sub-region also includes at least one of the following: The head position of the sub-region; The tail position of the sub-region; The length information of the text within the sub-region; Width information of the text within the sub-region; The height information of the text within the sub-region.

10. The method according to claim 8, wherein, The information of the selected text within the sub-region of the data structure, which is recorded in real time, and the corresponding region in the rendering area, is obtained through the following method: The starting and ending cursor positions of the selection operation are determined based on the cursor drag position when the user selects the text. Based on the cursor drag position, the starting cursor position, and the ending cursor position, information about the region portion of the selected text within the sub-region corresponding to the rendered area is obtained.

11. The method according to claim 7, wherein, Extracting the text selected by the selection operation from the plurality of components includes: Based on the information recorded in real time in the data structure corresponding to the sub-region where the selected text is located in the component selected by the selection operation, the text selected by the selection operation is extracted.

12. The method according to any one of claims 1-6, wherein, The acquisition of text from multiple components within the target area, and the text style information of the text, includes: Traverse the view tree of the target area of ​​the display interface, determine at least one view component within the target area, and obtain the text in multiple components of the target area and the text style information of the text from the at least one view component, wherein at least some of the multiple components are components of different types.

13. An electronic device, comprising: The processor, the communication interface, the memory, and the communication bus are provided, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The memory is used to store computer programs; The processor is configured to perform the method of any one of claims 1-12 by running the computer program stored in the memory.

14. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-12.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1-12.