Screen content identification method and device, program product and storage medium
This screen content recognition method, which allows users to draw a trajectory line on the screen to specify the recognition area, solves the problems of long recognition time and high power consumption in existing technologies, and enables flexible cross-line recognition and extraction, thereby improving the user experience.
Patent Information
- Application Number
- CN202511176814.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-12-30
AI Technical Summary
Current technologies that recognize all text in an image result in long recognition times and increased power consumption, and cannot enable users to select text across lines in an image.
A screen content recognition method is provided, which allows users to specify the areas to be recognized or not to be recognized by drawing trajectory lines on the screen, and uses polygon fitting technology to determine the specified markers to recognize only the target content required by the user.
It shortens recognition time, saves power consumption, and enables flexible cross-line recognition and extraction, thus improving the user experience.
Smart Images

Figure CN121236741A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed with the Chinese Patent Office with application number 202211473909.X, filing date November 22, 2022, entitled "Screen Content Recognition Method, Apparatus, Device and Storage Medium". Technical Field
[0002] This application relates to the field of computer technology, and in particular to a screen content recognition method, device, program product, and storage medium. Background Technology
[0003] With the development of computer technology, electronic devices such as mobile phones, tablets, and laptops have become an indispensable part of people's daily lives. Users can use electronic devices to browse content. When browsing images and other content on electronic devices, users may need to recognize the text portions they contain.
[0004] In related technologies, when recognizing text within an image, all text within the image is identified. However, recognizing the entire image not only leads to longer recognition times but also increases power consumption. Summary of the Invention
[0005] This application provides a screen content recognition method, apparatus, device, and storage medium, which can shorten recognition time and save power consumption. The technical solution is as follows:
[0006] Firstly, a screen content recognition method is provided. This method is applied to electronic devices. In this method, if a screen content recognition instruction is received, n trajectory lines drawn on the screen of the electronic device are obtained, where n is a positive integer. Then, a specified marker is obtained based on the n trajectory lines, and the target content is recognized from the screen content of the electronic device based on the specified marker.
[0007] Optionally, the screen content can be an image or interface displayed by the electronic device, which can be an application interface, a video playback interface, or a camera preview interface, etc., but is not limited to these. That is, this application can be applied to identify the content in various images or interfaces displayed by electronic devices, and can be applied to a variety of scenarios, making it convenient for users.
[0008] Optionally, the specified markers include a first marker and / or a second marker, wherein the first marker is used to indicate content that needs to be identified, and the second marker is used to indicate content that does not need to be identified.
[0009] As an example, the first marker is a closed shape, and the second marker is an open shape. In this case, the first marker is used to define the content to be identified, and the second marker is used to obscure the content that does not need to be identified. Optionally, the user can freely draw a trajectory line on the screen to ultimately draw the first and second markers.
[0010] As another example, the first mark is a first preset graphic, and the second mark is a second preset graphic, the first preset graphic and the second preset graphic have different shapes.
[0011] For example, the first preset shape is a closed shape, and the second preset shape is also a closed shape. The first preset shape can be a convex polygon, and the second preset shape can be a concave polygon, or the first and second preset shapes can be convex polygons of different shapes. In this case, the first marker is used to delineate the content to be identified, and the second marker is used to delineate the content that does not need to be identified.
[0012] For example, the first preset shape is a closed shape, and the second preset shape is an open shape. In this case, the first marker is used to define the content that needs to be identified, and the second marker is used to obscure the content that does not need to be identified.
[0013] Optionally, the electronic device may provide a first graphic option and a second graphic option. After selecting the first graphic option, the user can draw the trajectory line of a first preset graphic on the screen. After selecting the second graphic option, the user can draw the trajectory line of a second preset graphic on the screen.
[0014] As another example, both the first and second markers are closed shapes, but their lines are of different thicknesses. In this case, the first marker is used to define the content that needs to be identified, while the second marker is used to define the content that does not need to be identified.
[0015] Optionally, the electronic device can provide a first line option and a second line option. After selecting the first line option, the user can draw a line of a first thickness on the screen. After selecting the second preset graphic, the user can draw a line of a second thickness on the screen. The first thickness and the second thickness are different; the first thickness is the line thickness of the first marker, and the second thickness is the line thickness of the second marker.
[0016] As another example, both the first and second markers are closed shapes, but their lines are different colors. In this case, the first marker is used to define the content that needs to be identified, and the second marker is used to define the content that does not need to be identified.
[0017] Optionally, the electronic device can provide a first color option and a second color option. When the user selects the first color option, a trajectory line of the first color can be drawn on the screen. When the user selects the second color option, a trajectory line of the second color can be drawn on the screen. The first color and the second color are different; the first color is the color of the line marked with the first marker, and the second color is the color of the line marked with the second marker.
[0018] In this application, a user can draw a trajectory line on the screen. The electronic device can determine a specified marker (i.e., a first marker and / or a second marker) based on the trajectory line drawn by the user. Then, it can determine which content in the screen content needs to be recognized and which content does not need to be recognized based on the specified marker, and thus identify the target content that meets the user's needs from the screen content.
[0019] This application eliminates the need for full-text recognition of the electronic device's screen content. Instead, it identifies target content matching the user's needs from the screen content based on user-drawn markers. Recognition is quick, allowing for rapid content identification and saving power. Furthermore, since users can select which content to recognize or exclude—that is, they can select a local area or multiple discontinuous areas—this application enables cross-line recognition of screen content, making screen content recognition more flexible.
[0020] Optionally, the first marker is a closed shape, and the second marker is a non-closed shape. The operation of obtaining a specified marker based on the n trajectory lines can be as follows: perform polygon fitting processing on the n trajectory lines; if a polygon is fitted through at least one of the n trajectory lines, then the fitted polygon is determined as the first marker, and the trajectory lines among the n trajectory lines that do not fit polygons are determined as the second marker; if multiple polygons are fitted through at least one of the n trajectory lines, then the first marker is determined based on the overlap between the fitted polygons, and the trajectory lines among the n trajectory lines that do not fit polygons are determined as the second marker; if no polygon is fitted through any of the n trajectory lines, then the n trajectory lines are determined as the second marker.
[0021] Since the trajectory lines drawn freely by users are usually not very precise and regular, this application obtains more regular polygons by performing polygon fitting on the drawn trajectory lines, which can restore the user's drawing intention to a certain extent. This makes it easier to accurately determine the first and second marks, and thus accurately determine the user's recognition needs.
[0022] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to delineate the content not to be identified or to obscure the content not to be identified. The operation of identifying target content from the screen content of an electronic device based on the specified marker can be as follows: if the specified marker includes the first marker but does not include the second marker, then the target content is identified from the screen content of the electronic device based on the first area delineated by the first marker; if the specified marker includes the second marker but does not include the first marker, then the target content is identified from the screen content of the electronic device based on the second area delineated by the second marker or based on the content obscured by the second marker.
[0023] In this application, target content can be identified from the screen content of an electronic device based on a first defined area, or based on a second defined area or content obscured by a second mark. Thus, this application eliminates the need for full-text recognition; instead, specific target content can be selected for recognition based on user needs. Recognition time is short, content recognition can be completed quickly, and power consumption is saved.
[0024] Optionally, the operation of identifying target content from the screen content of an electronic device based on the first region defined by the first mark can include the following four cases:
[0025] The first scenario: If the content within the first area of the electronic device's screen is text, then the text paragraph containing that text is defined as the first paragraph. Based on the coordinates of each sentence in the first paragraph and the coordinate range of the first area, the overlap between each sentence in the first paragraph and the first area is determined. The target sentence is then identified from the first paragraph based on this overlap, and the target sentence is the target content. Alternatively, based on the coordinates of each character in the first paragraph and the coordinate range of the first area, the overlap between each character in the first paragraph and the first area is determined. The target character is then identified from the first paragraph based on this overlap, and the target character is the target content.
[0026] In this application, the actual outlined area (i.e., the first area) can be combined with a text paragraph. Then, the overlap between the sentence or character coordinates of the text paragraph and the outlined area can be analyzed using the sentence or character coordinates of the text paragraph, thereby determining the target sentence or target character to be identified. In this way, even if the user does not outline the entire sentence or only half a character, the user's outlined intention can still be restored to a certain extent, thus accurately identifying the target sentence or target character that meets the user's needs.
[0027] The second scenario: If the content in the first area of the electronic device's screen is table content, then the table to which the table content in the first area belongs is determined to be the first table; based on the coordinates of each cell in the first table and the coordinate range of the first area, the overlap between each cell in the first table and the first area is determined; based on the overlap between each cell in the first table and the first area, the text in the target cell is identified from the first table, and the text in the target cell is the target content.
[0028] In this application, the actual outlined area (i.e., the first area) can be combined with a table. Then, the overlap between the table cells and the outlined area can be analyzed using the table's cell coordinates to determine the target cell to be identified. In this way, even if the user does not completely outline the cell, the user's outlined intention can be reproduced to a certain extent, thereby accurately identifying the text in the target cell that meets the user's needs.
[0029] The third scenario: If the content located in the first area of the electronic device's screen is an image, then the image to which the image content in the first area belongs is determined to be the first image; the first image is identified based on its coordinates, and the first image is the target content.
[0030] This application allows users to select images. Users can identify images by drawing a first mark on them, thus improving the flexibility of screen content recognition.
[0031] The fourth scenario: If the content in the first area of the electronic device's screen is formula content, then the formula containing the formula content in the first area is identified, and the formula containing the formula content in the first area is the target content.
[0032] This application allows users to select formulas. Users can identify formulas by marking the first mark on them, thus improving the flexibility of screen content recognition.
[0033] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to delineate content that does not need to be identified or to obscure content that does not need to be identified. The operation of identifying target content from the screen content of an electronic device based on the specified markers can be as follows: If the specified markers include both the first and second markers, and the first and second markers do not overlap, then the first marker is retained and the second marker is discarded from the specified markers, and the target content is identified from the screen content of the electronic device based on the first area delineated by the first marker. Alternatively, if the specified markers include both the first and second markers, and the first and second markers do not overlap, then the first marker is discarded and the second marker is retained from the specified markers, and the target content is identified from the screen content of the electronic device based on the second area delineated by the second marker or based on the content obscured by the second marker.
[0034] In this application, if both the first marker and the second marker exist but do not overlap, then their delineation intentions conflict. In this case, one of them can be discarded while the other is retained so that content recognition can continue normally.
[0035] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to delineate the content that does not need to be identified. The operation of identifying target content from the screen content of an electronic device based on the specified markers can be as follows: if the specified markers include the first marker and the second marker, and the second marker is located within the first marker, then the area outside the second area delineated by the first marker in the first region is determined as the third region, and the target content is identified from the screen content of the electronic device based on the third region.
[0036] In this application, if the second mark is located within the first mark, it means that the user wants to identify the content of other areas (i.e., the third area) in the first area enclosed by the first mark but excluding the second area enclosed by the second mark. Therefore, the target content can be identified from the screen content of the electronic device based on the third area, and the identified target content meets the user's needs.
[0037] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to obscure content that does not need to be identified. The operation of identifying target content from the screen content of an electronic device based on the specified marker can be as follows: if the specified marker includes the first marker and the second marker, and the second marker is located within the first marker, then the target content is identified from the screen content of the electronic device based on the first area delineated by the first marker, and the content obscured by the second marker is deleted from the identified target content.
[0038] In this application, if the second mark is located within the first mark, it means that the user wants to identify the content in the first area enclosed by the first mark, excluding the content obscured by the second mark. Therefore, the target content can be identified from the screen content of the electronic device based on the first area enclosed by the first mark, and then the content obscured by the second mark can be deleted from the identified target content. The target content obtained in this way meets the user's needs.
[0039] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to delineate the content that does not need to be identified. The operation of identifying target content from the screen content of an electronic device based on the specified markers can be as follows: if the specified markers include the first marker and the second marker, and the first marker intersects with the second marker, then the first marker is retained and the intersection of the first marker and the second marker is used as the new second marker; or, the second marker is retained and the union of the second marker and the first marker is used as the new first marker. The area outside the second area delineated by the second marker within the first region defined by the first marker is determined as the third region, and the target content is identified from the screen content of the electronic device based on the third region.
[0040] In this application, if the first mark and the second mark intersect, it means that the user is likely to want to identify the content of the area outside the intersection of the first mark and the second mark (i.e., the third area) in the first area enclosed by the first mark. Therefore, the target content can be identified from the screen content of the electronic device based on the third area, and the identified target content meets the user's needs.
[0041] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to obscure content that does not need to be identified. The operation of identifying target content from the screen content of an electronic device based on the specified markers can be as follows: if the specified markers include a first marker and a second marker, and the first marker intersects with the second marker, then the first marker is retained, and the portion of the second marker located within the first marker is taken as a new second marker; the target content is identified from the screen content of the electronic device based on the first area delineated by the first marker, and the content obscured by the second marker is deleted from the identified target content.
[0042] In this application, if the first mark and the second mark intersect, it means that the user is likely to want to identify the content in the first area enclosed by the first mark, excluding the content partially obscured by the second mark located within the first mark. Therefore, the target content can be identified from the screen content of the electronic device based on the first area enclosed by the first mark, and then the partially obscured content of the second mark located within the first mark can be deleted from the identified target content. The target content obtained in this way meets the user's needs.
[0043] Furthermore, after identifying target content from the screen content of the electronic device based on specified markers, the identified target content can be extracted. Specifically, the identified target content can be stored in the system clipboard. Subsequently, if any edit box is detected to have focus, the target content stored in the system clipboard is pasted into that edit box.
[0044] This application enables cross-line recognition of screen content on electronic devices. Therefore, after extracting the identified target content, cross-line extraction of screen content is achieved, thus improving the flexibility of screen content extraction. Furthermore, since the edit box with focus is highly likely to be the edit box the user currently needs to use, the identified target content can be directly pasted into the edit box with focus in this embodiment, thereby facilitating user use and improving the user experience.
[0045] Secondly, a screen content recognition device is provided, which has the function of implementing the screen content recognition method described in the first aspect. The screen content recognition device includes at least one module, which is used to implement the screen content recognition method provided in the first aspect.
[0046] Thirdly, a screen content recognition device is provided, comprising a processor and a memory. The memory stores a program that supports the screen content recognition device in executing the screen content recognition method provided in the first aspect, and stores data related to implementing the screen content recognition method described in the first aspect. The processor is configured to execute the program stored in the memory. The screen content recognition device may further include a communication bus for establishing a connection between the processor and the memory.
[0047] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the screen content recognition method described in the first aspect.
[0048] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the screen content recognition method described in the first aspect.
[0049] The technical effects achieved by the second, third, fourth, and fifth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0051] Figure 2 This is a block diagram of a software system for an electronic device provided in an embodiment of this application;
[0052] Figure 3 This is a schematic diagram of an image display provided in an embodiment of this application;
[0053] Figure 4 This is a schematic diagram of a trajectory line provided in an embodiment of this application;
[0054] Figure 5 This is a schematic diagram of an application interface provided in an embodiment of this application;
[0055] Figure 6 This is a schematic diagram of a video playback interface provided in an embodiment of this application;
[0056] Figure 7 This is a schematic diagram of a camera preview interface provided in an embodiment of this application;
[0057] Figure 8 This is a schematic diagram of the first type of first mark and second mark provided in the embodiments of this application;
[0058] Figure 9 This is a schematic diagram of a first preset graphic provided in an embodiment of this application;
[0059] Figure 10 This is a schematic diagram of a second preset graphic provided in an embodiment of this application;
[0060] Figure 11 This is a schematic diagram of the second type of first mark and second mark provided in the embodiments of this application;
[0061] Figure 12 This is a flowchart of a screen content recognition method provided in an embodiment of this application;
[0062] Figure 13 This is a schematic diagram of a notification bar provided in an embodiment of this application;
[0063] Figure 14 This is a schematic diagram of a graphical option provided in an embodiment of this application;
[0064] Figure 15 This is a schematic diagram of polygon fitting and merging provided in an embodiment of this application;
[0065] Figure 16 This is a schematic diagram of another trajectory line provided in an embodiment of this application;
[0066] Figure 17 This is a schematic diagram of a first region provided in an embodiment of this application;
[0067] Figure 18 This is a schematic diagram of an edit box provided in an embodiment of this application;
[0068] Figure 19 This is a schematic diagram of a table provided in an embodiment of this application;
[0069] Figure 20 This is a schematic diagram of an image correction method provided in an embodiment of this application;
[0070] Figure 21 This is a schematic diagram illustrating formula recognition provided in an embodiment of this application;
[0071] Figure 22 This is a schematic diagram of another table provided in an embodiment of this application;
[0072] Figure 23 This is a flowchart illustrating the processing of the first and second marks according to the embodiments of this application;
[0073] Figure 24 This is a schematic diagram illustrating the processing of the first and second marks provided in the embodiments of this application;
[0074] Figure 25 This is a flowchart illustrating the processing of the second type of first and second marks provided in the embodiments of this application;
[0075] Figure 26 This is a schematic diagram illustrating the processing of the second type of first and second marks provided in the embodiments of this application;
[0076] Figure 27 This is a flowchart illustrating the third type of processing of the first and second marks provided in the embodiments of this application;
[0077] Figure 28 This is a schematic diagram illustrating the processing of the third type of first and second marks provided in the embodiments of this application;
[0078] Figure 29 This is a flowchart illustrating the processing of the fourth type of first and second markers provided in the embodiments of this application;
[0079] Figure 30 This is a schematic diagram illustrating the processing of the fourth type of first and second marks provided in the embodiments of this application;
[0080] Figure 31 This is a flowchart illustrating the processing of the fifth type of first and second markers provided in the embodiments of this application;
[0081] Figure 32 This is a schematic diagram illustrating the processing of the fifth type of first and second marks provided in the embodiments of this application;
[0082] Figure 33 This is a flowchart illustrating video content recognition provided in an embodiment of this application;
[0083] Figure 34 This is a schematic diagram of the third type of first and second markings provided in the embodiments of this application;
[0084] Figure 35 This is a schematic diagram of a screen content recognition method provided in an embodiment of this application;
[0085] Figure 36 This is a flowchart of another screen content recognition method provided in an embodiment of this application;
[0086] Figure 37 This is a schematic diagram of the first content recognition processing procedure provided in the embodiments of this application;
[0087] Figure 38 This is a schematic diagram of the second content recognition processing procedure provided in the embodiments of this application;
[0088] Figure 39 This is a schematic diagram of the third content recognition processing procedure provided in the embodiments of this application;
[0089] Figure 40 This is a schematic diagram of the fourth content recognition processing procedure provided in the embodiments of this application;
[0090] Figure 41 This is a schematic diagram of the structure of a screen content recognition device provided in an embodiment of this application. Detailed Implementation
[0091] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0092] It should be understood that "multiple" as mentioned in this application refers to two or more. In the description of this application, unless otherwise stated, " / " indicates "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, to facilitate a clear description of the technical solutions of this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and that "first," "second," etc., do not necessarily imply differences.
[0093] The terms "one embodiment" or "some embodiments" used in this application mean that one or more embodiments of this application include the specific features, structures, or characteristics described in that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this application do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. Furthermore, the terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0094] The electronic devices involved in the embodiments of this application will be described below.
[0095] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. See also... Figure 1The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a SIM card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0096] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0097] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0098] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0099] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0100] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device 100 via the power management module 141.
[0101] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0102] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0103] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0104] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0105] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0106] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0107] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions, such as saving music, video, and other files on the external memory card.
[0108] Internal memory 121 can be used to store computer-executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created by electronic device 100 during use (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0109] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D and application processor.
[0110] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0111] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is an integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0112] The software system of electronic device 100 will be described next.
[0113] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses a layered Android system as an example to illustrate the software system of electronic device 100.
[0114] Figure 2 This is a block diagram of a software system for an electronic device 100 provided in an embodiment of this application. See also... Figure 2 A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system includes an application layer, an application framework layer, the Android runtime, a system layer, and a kernel layer.
[0115] The application layer can include a series of applications. For example... Figure 2 As shown, applications can include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, SMS, and other applications.
[0116] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 2As shown, the application framework layer can include a window manager, content providers, a view system, a phone manager, a resource manager, and a notification manager. The window manager manages window programs. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture the screen. The content provider stores and retrieves data, making this data accessible to the application. This data can include videos, images, audio, made and received phone calls, browsing history and bookmarks, and phone books. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build the application's display interface, which can consist of one or more views, such as a view displaying SMS notification icons, a view displaying text, and a view displaying images. The phone manager provides communication functions for the electronic device 100, such as managing call status (including connection and disconnection). The resource manager provides the application with various resources, such as localized strings, icons, images, layout files, and video files. The notification manager allows the application to display notification information in the status bar, which can be used to convey informational messages and can disappear automatically after a short pause without user interaction. For example, the notification manager is used to notify users of download completions and message alerts. The notification manager can also display notifications as icons or scrolling text in the system's top status bar, such as notifications from background applications. Furthermore, the notification manager can appear as dialog boxes on the screen, such as displaying text messages in the status bar, emitting sounds, causing electronic devices to vibrate, or flashing indicator lights.
[0117] The Android Runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core libraries consist of two parts: one part contains the functionalities that Java needs to call, and the other part is the core Android library itself. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0118] The system library can include multiple functional modules, such as: a surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), and 2D graphics engines (e.g., SGL). The surface manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The media libraries support playback and recording of various common audio and video formats, as well as still image files. The media libraries support various audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. The 3D graphics processing libraries are used to implement 3D graphics drawing, image rendering, compositing, and layer processing. The 2D graphics engine is the drawing engine for 2D graphics.
[0119] The kernel layer is the layer between hardware and software. It contains display drivers, camera drivers, audio drivers, sensor drivers, and more.
[0120] The application scenarios involved in the embodiments of this application are described below.
[0121] With the development of computer technology, electronic devices such as mobile phones, tablets, and laptops have become an indispensable part of people's daily lives. Users can use electronic devices to browse content. In some scenarios, such as... Figure 3 As shown, when users browse images and other content on electronic devices, they may need to recognize and extract the text contained therein.
[0122] In related technologies, when recognizing and extracting text from an image, all text within the image is recognized and extracted. Specifically, all text within the image is first recognized, and then a text extraction icon is displayed. After clicking the text extraction icon, the user can perform extraction operations such as copying, translating, sharing, and searching on the recognized text.
[0123] However, the above methods have several drawbacks. First, recognizing the entire text results in a long recognition time; considering the amount of text and computing power, the recognition time can even reach several seconds, leading to a delay in the display of the extracted text icons. Second, recognizing the entire text also increases power consumption. Furthermore, since the final recognized text is the entire text, and users can only select text consecutively, it is impossible for users to select text across multiple lines within an image.
[0124] Therefore, embodiments of this application provide a screen content recognition method that allows users to flexibly select and / or not recognize content within the screen. For example, such as Figure 4 As shown in Figure (a), a user browses images on a mobile phone. In this case, as... Figure 4As shown in Figure (b), the user can draw a trajectory line on the phone screen. This trajectory line can define region A in the image, instructing the user to select and recognize content within region A, while ignoring content in other regions. Or, as... Figure 4 As shown in Figure (c), the user can draw a trajectory line on the phone screen. This trajectory line can define region B in the image, indicating to the user that they should not recognize the content in region B, but instead recognize the content in other regions. Thus, in this embodiment, it is not necessary to recognize the entire image; instead, the user can select specific regions to recognize, resulting in a short recognition time, rapid content recognition, and energy savings. Furthermore, since the user can choose which regions to recognize or not recognize based on their needs, cross-line recognition of screen content can be achieved. This allows the user to select recognized content across multiple lines, making screen content recognition and extraction more flexible.
[0125] The screen content involved in the embodiments of this application will be described below.
[0126] The screen content recognition method provided in this application can be used to recognize the screen content of an electronic device. The electronic device can be a mobile phone, tablet computer, laptop computer, television set, conference equipment, etc., and this application does not limit this to any particular type. The screen content refers to the content displayed by the electronic device.
[0127] Optionally, such as Figure 3 As shown, the screen content can be an image displayed by an electronic device. In this case, the screen content recognition method provided in this application embodiment can be used to identify the content contained in the image displayed by the electronic device.
[0128] Optionally, the screen content can be the interface displayed by the electronic device, such as an application interface, a video playback interface, a camera preview interface, etc., but not limited to these. In this case, the screen content recognition method provided in the embodiments of this application can be used to identify the content contained in the interface displayed by the electronic device.
[0129] For example, such as Figure 5 As shown, the screen content can be the application interface 501 of an information application displayed on an electronic device, which is used to display information. In this case, the screen content recognition method provided in this application embodiment can be used to identify the information displayed in the application interface 501 of the information application.
[0130] For example, such as Figure 6As shown, the screen content can be a video playback interface 601 of a video application displayed on an electronic device, which is used to play videos. Before executing the screen content recognition method provided in this application embodiment, the video playing in the video playback interface 601 of the video application may be in a paused state. In this case, the screen content recognition method provided in this application embodiment can be used to recognize the video image of the paused video in the video playback interface 601 of the video application.
[0131] For example, such as Figure 7 As shown, the screen content can be the camera preview interface 701 of a camera application displayed on an electronic device. This camera preview interface 701 displays a preview image captured by the camera. In this case, the screen content recognition method provided in this application embodiment can be used to identify the preview image displayed in the camera preview interface 701 of the camera application.
[0132] Optionally, the screen content recognition method provided in this application embodiment is applicable to recognizing text in various fonts contained in the screen content. For example, it can recognize text in fonts such as printed, handwritten, and calligraphic fonts contained in the screen content. This application embodiment does not limit this.
[0133] The designated markings involved in the embodiments of this application will be described below.
[0134] The screen content recognition method provided in this application can determine the content to be recognized and / or the content that does not need to be recognized based on specified marks drawn by the user on the screen of an electronic device when recognizing the screen content. The specified marks involved in this application include a first mark and / or a second mark. The first mark is used to indicate the content to be recognized, and the second mark is used to indicate the content that does not need to be recognized.
[0135] The first and second markers are different so that the electronic device can distinguish between them and identify target content that meets the user's needs from the screen content based on the first and second markers. Several possible implementations of the first and second markers are described below:
[0136] In one possible implementation, the first marker can be a closed shape, while the second marker can be an open shape. In this case, the first marker is used to define the content to be identified, and the second marker is used to obscure the content that does not need to be identified.
[0137] Optionally, users can freely draw trajectory lines on the screen to ultimately draw the first and second marks.
[0138] For example, such as Figure 8As shown in Figure (a), the first marker can be a closed shape formed by the trajectory line A freely drawn by the user, and the content enclosed by the first marker is the content to be identified. Figure 8 As shown in Figure (b), the second marker can be a non-closed shape formed by the trajectory line B freely drawn by the user, and the content covered by the second marker is the content that does not need to be recognized.
[0139] In a second possible implementation, the first marker can be a predefined first preset graphic, and the second marker can be a predefined second preset graphic. The first preset graphic and the second preset graphic have different shapes.
[0140] As an example, the first preset shape is a closed shape, and the second preset shape is also a closed shape. For instance, the first preset shape can be a convex polygon, and the second preset shape can be a concave polygon; or, the first preset shape and the second preset shape can be convex polygons of different shapes. In this case, the first marker is used to delineate the content that needs to be identified, and the second marker is used to delineate the content that does not need to be identified.
[0141] As another example, the first preset shape is a closed shape, and the second preset shape is an open shape. In this case, the first marker is used to define the content that needs to be identified, and the second marker is used to obscure the content that does not need to be identified.
[0142] Optionally, the electronic device can provide multiple graphic options. Users can select different graphic options to draw the trajectory lines of different graphics on the screen, ultimately drawing a first mark and a second mark. For example, the electronic device can provide a first graphic option and a second graphic option. After selecting the first graphic option, the user can draw the trajectory line of a first preset graphic on the screen. After selecting the second graphic option, the user can draw the trajectory line of a second preset graphic on the screen.
[0143] In some embodiments, the first preset graphic may include a plurality of predefined graphics. In this case, the electronic device may provide a plurality of first graphic options, each corresponding one-to-one with one of the plurality of graphics. After a user selects one of the plurality of first graphic options, the user can draw the trajectory line of the graphic corresponding to that first graphic option on the screen.
[0144] In some embodiments, the second preset graphic may include a plurality of predefined graphics. In this case, the electronic device may provide multiple second graphic options, each corresponding to one of the multiple graphics. After the user selects one of the multiple second graphic options, the user can draw the trajectory line of the graphic corresponding to that second graphic option on the screen.
[0145] For example, the first preset graphic may include predefined shapes. Figure 9 The second preset figure may include at least one of the three closed figures shown in Figures (a), (b), and (c), but is not limited to these. Figure 10 At least one of the three non-closed figures shown in Figures (a), (b), and (c), but not limited to these.
[0146] In a third possible implementation, both the first and second markers can be closed shapes, but their line thicknesses differ. The line thickness of both markers can be preset. In this case, the first marker is used to delineate the content to be recognized, and the second marker is used to delineate the content that does not need to be recognized.
[0147] Optionally, the electronic device can provide two line options. Users can freely draw lines of varying thicknesses on the screen by selecting different line options to ultimately create the first and second marks. For example, the electronic device can provide a first line option and a second line option. After selecting the first line option, the user can draw a line of a first thickness on the screen. After selecting the second preset shape, the user can draw a line of a second thickness on the screen. The first and second thicknesses are different; the first thickness is the line thickness of the first mark, and the second thickness is the line thickness of the second mark.
[0148] For example, such as Figure 11 As shown in Figure (a), the first marker can be a closed shape formed by the trajectory line A freely drawn by the user, and the content enclosed by the first marker is the content to be identified. Figure 11 As shown in Figure (b), the second marker can be a closed shape formed by the trajectory line B freely drawn by the user. The content enclosed by the second marker is the content that does not need to be recognized. The line thickness of trajectory line B is different from that of trajectory line A.
[0149] A fourth possible implementation involves both the first and second markers being closed shapes, but with different line colors. The line colors of both markers can be preset. In this case, the first marker is used to define the content to be identified, while the second marker is used to define the content that does not need to be identified.
[0150] Optionally, the electronic device can provide two color options. Users can freely draw lines of different colors on the screen by selecting different color options to ultimately draw the first and second marks. For example, the electronic device can provide a first color option and a second color option. After selecting the first color option, the user can draw a line of the first color on the screen. After selecting the second color option, the user can draw a line of the second color on the screen. The first and second colors are different; the first color is the line color of the first mark, and the second color is the line color of the second mark.
[0151] For example, the first marker can be green, and the second marker can be red. In this case, the first marker can be a closed shape formed by a green line drawn freely by the user, and the content enclosed by the first marker is the content that needs to be identified. The second marker can be a closed shape formed by a red line drawn freely by the user, and the content enclosed by the second marker is the content that does not need to be identified.
[0152] It should be noted that in this embodiment of the application, the user can draw a trajectory line on the screen. The electronic device can determine the specified marker (i.e., the first marker and / or the second marker) based on the trajectory line drawn by the user. Then, it can determine which content in the screen content needs to be identified and which content does not need to be identified based on the specified marker. In this way, the target content that meets the user's needs can be identified from the screen content.
[0153] The screen content recognition method provided in the embodiments of this application will be explained in detail below.
[0154] Figure 12 This is a flowchart illustrating a screen content recognition method provided in an embodiment of this application. This method can be applied to electronic devices, which can be the aforementioned... Figures 1 to 2 The electronic device described in the embodiments. See also Figure 12 The method includes the following steps:
[0155] Step 1201: If the electronic device receives a screen content recognition instruction, the electronic device acquires the n trajectory lines drawn on the screen of the electronic device, where n is a positive integer.
[0156] Screen content recognition commands are used to instruct the recognition of screen content on an electronic device. These commands can be triggered by the user within the device. Screen content recognition commands can be triggered in several ways. In one possible way, the electronic device determines that it has received a screen content recognition command if it detects a selection of a screen recognition button. For example, ... Figure 13As shown, users can swipe down on the screen of an electronic device to open the notification bar 1301. The notification bar 1301 can display buttons such as WLAN, mobile data, mute, auto-rotate, Bluetooth, and screen recognition 1302. After the user clicks the screen recognition button, a screen content recognition command is triggered on the electronic device, which then receives the command. The user can then draw a trajectory line on the screen, and the electronic device can capture the trajectory line drawn by the user.
[0157] In the first possible implementation, if the difference between the first and second markers is that the first marker is a closed shape while the second marker is a non-closed shape, for example, the first and second markers are based on... Figure 8 As shown, this is achieved by allowing the user to freely draw lines on the screen of an electronic device after triggering a screen content recognition command. This allows the user to create closed and / or open shapes. The electronic device can capture the lines drawn by the user on the screen.
[0158] In the second possible implementation, if the difference between the first mark and the second mark is that the first mark is a first preset graphic, and the second mark is a second preset graphic, for example, the first mark is based on... Figure 9 It is implemented in the manner shown, the second marker is... Figure 10 As shown, after a user triggers a screen content recognition command on an electronic device, the device can display a first graphic option and a second graphic option. In this case, if the user selects the first graphic option, they can draw the trajectory of a first preset graphic on the screen; if the user selects the second graphic option, they can draw the trajectory of a second preset graphic on the screen. During the process of the user drawing the trajectory on the screen, the electronic device can acquire the trajectory drawn by the user.
[0159] For example, after a user clicks the selected screen recognition button on an electronic device, such as Figure 14 As shown in Figure (a), the electronic device can display a first graphics option 1401 and a second graphics option 1402 floating above the current screen content. Then, as... Figure 14 As shown in Figure (b), the user can click on the first graphic option 1401, and then the user can draw the trajectory line A of the first preset graphic on the screen of the electronic device. Figure 14 As shown in Figure (c), the user can click on the second graphic option 1402, and then the user can draw the trajectory line B of the second preset graphic on the screen of the electronic device.
[0160] In a third possible implementation, if the difference between the first and second markers is that both the first and second markers are closed shapes, but their line thicknesses differ—for example, the first and second markers are based on different line thicknesses—then... Figure 11 As shown, this is achieved by having the user trigger a screen content recognition command on the electronic device, which then displays a first line option and a second line option. In this case, selecting the first line option allows the user to draw a line of a first thickness on the screen, while selecting the second line option allows them to draw a line of a second thickness. The first and second thicknesses are different; the first thickness is the thickness of the line marked with a first marker, and the second thickness is the thickness of the line marked with a second marker. During the process of the user drawing lines on the screen, the electronic device can capture the lines drawn by the user.
[0161] In a fourth possible implementation, if the difference between the first and second markers is that both are closed shapes, but their line colors differ, then after the user triggers a screen content recognition command on the electronic device, the device can display a first color option and a second color option. For example, after the user clicks the screen recognition button, the device can display the first and second color options floating above the current screen content. In this case, if the user selects the first color option, they can draw a trajectory line of the first color on the screen; if they select the second color option, they can draw a trajectory line of the second color. The first and second colors are different; the first color is the line color of the first marker, and the second color is the line color of the second marker. During the process of the user drawing the trajectory line on the screen, the electronic device can acquire the trajectory line drawn by the user.
[0162] Optionally, when a user draws lines on the screen of an electronic device, after drawing several lines, they can adjust the size of the area enclosed by the drawn lines or the size of the area they obscure by dragging. For example, after drawing the line of a closed shape on the screen, the user can adjust the size of the closed shape by dragging, thus adjusting the size of the area enclosed by the closed shape. Similarly, after drawing the line of a non-closed shape on the screen, the user can adjust the size of the non-closed shape by dragging, thus adjusting the size of the area obscured by the non-closed shape.
[0163] As an example, the electronic device can also provide an "End Recognition" button. After completing the drawing, the user can click the "End Recognition" button to trigger an "End Recognition" command, at which point the selection of the screen recognition button will be deselected. After receiving the "End Recognition" command, the electronic device can obtain a specified marker (including a first marker and / or a second marker) based on all the trajectory lines (i.e., n trajectory lines) drawn by the user on the screen of the electronic device, and then execute step 1202 below. For example, the "End Recognition" button can be displayed in the drop-down notification bar of the electronic device, or the "End Recognition" button can be displayed as a floating button on top of the screen content of the electronic device. Of course, the "End Recognition" button can also be displayed in other forms, which are not limited in this embodiment.
[0164] As another example, after completing the drawing, the user can click the screen recognition button again to deselect it, which will also trigger the end recognition command. After receiving the end recognition command, the electronic device can obtain the specified marker based on all the trajectory lines drawn by the user on the screen of the electronic device, that is, it can perform the following step 1202.
[0165] As another example, after a preset time has elapsed since the user clicked the selected screen recognition button, the electronic device can automatically deselect the screen recognition button, thus triggering an end-of-recognition command. Upon receiving the end-of-recognition command, the electronic device can obtain the specified marker based on all the trajectory lines drawn by the user on the screen, and can then execute step 1202 below.
[0166] It should be noted that, in this embodiment, the user can draw a trajectory line on the screen of the electronic device in various ways. For example, the user can draw a trajectory line on the screen of the electronic device by touching the screen with their finger, or by using a stylus, or by using a mouse, keyboard, or by using voice operation, motion control, or air gesture operation. Of course, the user can also draw a trajectory line on the screen of the electronic device in other ways, and this embodiment does not limit this method.
[0167] Optionally, the trajectory lines drawn by the user on the screen of the electronic device can be displayed on the screen. That is, after acquiring the trajectory lines drawn by the user on the screen, the electronic device can display the trajectory lines on the screen. In this case, after receiving the end recognition command, the electronic device can also perform a screenshot operation to obtain a screenshot, which includes the screen content displayed by the electronic device and the n trajectory lines drawn by the user on the screen content. In this case, the electronic device can save the screenshot, and further, the electronic device can also remind the user to share the screenshot. The user can share the screenshot with others through the electronic device.
[0168] Step 1202: The electronic device obtains the specified marker based on the n trajectory lines.
[0169] Users draw lines on the screen of an electronic device to indicate content that needs to be recognized and / or content that does not need to be recognized. Therefore, after the electronic device obtains n lines drawn by the user on the screen, it can obtain a specified marker based on the n lines so that it can subsequently determine which content on the screen needs to be recognized and which does not.
[0170] In the first possible implementation, if the difference between the first mark and the second mark is that the first mark is a closed shape and the second mark is a non-closed shape, then the operation of step 1202 may include the following steps (1) to (4):
[0171] (1) The electronic device performs polygon fitting on the n trajectory lines.
[0172] Polygon fitting refers to approximating curves with straight lines and fitting irregular curves with polygons, thus fitting the curves into regular polygons. In the embodiments of this application, performing polygon fitting processing on the n trajectory lines involves attempting to fit one or more polygons using one or more of the n trajectory lines.
[0173] It is worth noting that since the trajectory lines drawn freely by users are usually not very precise and regular, the embodiments of this application obtain relatively regular polygons by performing polygon fitting on the drawn trajectory lines, which can restore the user's drawing intention to a certain extent. This makes it easier to accurately determine the first mark and the second mark, and thus accurately determine the user's recognition needs.
[0174] (2) If a polygon is fitted by at least one of the n trajectory lines, the electronic device will identify the fitted polygon as the first mark and the trajectory lines of the n trajectory lines that do not fit a polygon as the second mark.
[0175] If a polygon is fitted by at least one of the n trajectory lines, it means that the at least one trajectory line forms a closed figure (i.e., the polygon), and therefore the polygon formed can be identified as the first marker.
[0176] In this case, the other trajectory lines among the n trajectory lines, excluding the at least one trajectory line, are the trajectory lines that have not fitted the polygon. These trajectory lines that have not fitted the polygon form a non-closed figure, and therefore these trajectory lines that have not fitted the polygon can be identified as the second marker.
[0177] In particular, if all n trajectory lines fit a polygon, that is, if there are no trajectory lines among the n trajectory lines that do not fit a polygon, then there is no need to determine the second mark. In this case, the specified mark obtained by the electronic device based on the n trajectory lines only includes the first mark.
[0178] (3) If multiple polygons are fitted through at least one of the n trajectory lines, the electronic device determines the first mark based on the fitted multiple polygons and the overlap of the multiple polygons, and determines the trajectory line of the n trajectory lines that does not fit a polygon as the second mark.
[0179] The operation of the electronic device to determine the first mark based on the fitted polygons and their overlap can be as follows: if none of the polygons overlap, then the polygons are determined as the first mark. If there are overlapping polygons among the polygons, then the overlapping polygons are merged to obtain a new polygon, and the merged new polygon and the polygons among the polygons that do not overlap with other polygons are determined as the first mark.
[0180] If multiple polygons are fitted through at least one of the n trajectory lines, it means that the at least one trajectory line forms multiple closed figures (i.e., the multiple polygons). Therefore, the first marker can be determined based on the overlap of these multiple polygons. Specifically, for any two polygons among these multiple polygons, since both polygons are used to indicate the content to be identified, if these two polygons overlap, they can be merged to obtain a new polygon as the first marker. The area enclosed by the new polygon is the union of the areas enclosed by the two polygons.
[0181] In this case, the other trajectory lines among the n trajectory lines, excluding the at least one trajectory line, are the trajectory lines that have not fitted the polygon. These trajectory lines that have not fitted the polygon form a non-closed figure, and therefore these trajectory lines that have not fitted the polygon can be identified as the second marker.
[0182] In particular, if all n trajectory lines fit a polygon, that is, if there are no trajectory lines among the n trajectory lines that do not fit a polygon, then there is no need to determine the second mark. In this case, the specified mark obtained by the electronic device based on the n trajectory lines only includes the first mark.
[0183] (4) If no polygon is fitted through the n trajectory lines, the electronic device determines the n trajectory lines as the second marker.
[0184] If no polygon is fitted by any of the n trajectory lines, it means that the n trajectory lines form non-closed shapes. Therefore, the n trajectory lines can be identified as the second marker. In this case, the designated marker obtained by the electronic device based on the n trajectory lines only includes the second marker.
[0185] For example, such as Figure 15 As shown, the electronic device first performs polygon fitting on the n trajectory lines, and fits multiple polygons through at least one of the n trajectory lines, for example, as... Figure 16 As shown in Figure (a), the electronic device acquires trajectory lines 1, 2, 3, and 4 drawn on its screen. After performing polygon fitting on these four trajectory lines, the electronic device fits trajectory line 1 to obtain... Figure 16 Polygon 1 shown in Figure (b) is fitted with trajectory line 2. Figure 16 Polygon 2 shown in Figure (b) is fitted with trajectory line 3. Figure 16 In Figure (b), polygon 3 is shown, while trajectory line 4 does not fit a polygon. Subsequently, the electronic device merges the multiple fitted polygons based on their overlap to obtain a new polygon, for example, Figure 16 In Figure (b), polygon 1 overlaps with polygon 2. Therefore, polygon 1 and polygon 2 can be merged to obtain... Figure 16 The new polygon 4 is shown in Figure (c). The electronic device then identifies the newly merged polygon and the polygons that do not overlap with other polygons as the first marker, and the trajectory lines from the n trajectory lines that do not fit a polygon as the second marker. For example, it can... Figure 16 In Figure (c), the new polygons 4 and 3 are identified as the first markers, and the trajectory line 4 is identified as the second marker.
[0186] In the second possible implementation, if the difference between the first mark and the second mark is that the first mark is a first preset graphic and the second mark is a second preset graphic, then the operation of step 1202 can be as follows: for any one of the n trajectory lines, if the graphic formed by this trajectory line is the first preset graphic, then the electronic device determines the graphic formed by this trajectory line as the first mark; if the graphic formed by this trajectory line is the second preset graphic, then the electronic device determines the graphic formed by this trajectory line as the second mark.
[0187] In this implementation, each of the n trajectory lines drawn by the user on the screen of the electronic device is directly a trajectory line of either the first preset graphic or the second preset graphic. That is, the graphic formed by each of the n trajectory lines is directly the first preset graphic or the second preset graphic. Therefore, the electronic device can directly determine the first and second marks from the graphic formed by the n trajectory lines. In this case, the process of the electronic device acquiring the first and second marks is relatively simple and can improve acquisition efficiency.
[0188] In the third possible implementation, if the difference between the first mark and the second mark is that the first mark is a closed shape and the second mark is also a closed shape, but the line thickness of the first mark and the second mark is different, then the operation of step 1202 may include the following steps (1) to (3):
[0189] (1) The electronic device groups the n trajectory lines to obtain the first group of trajectory lines and / or the second group of trajectory lines.
[0190] The thickness of the lines in the first group of trajectory lines is the same as the thickness of the lines marked with the first mark, and the thickness of the lines in the second group of trajectory lines is the same as the thickness of the lines marked with the second mark.
[0191] Specifically, if the line thickness of each of the n trajectory lines is the same as the line thickness of the first marker, then after the electronic device groups the n trajectory lines, it will only obtain the first group of trajectory lines. In this case, there is no second marker, and the specified marker obtained by the electronic device based on the n trajectory lines only includes the first marker.
[0192] If the thickness of each of the n trajectory lines is the same as the thickness of the second marker, then after the electronic device groups the n trajectory lines, it will only obtain the second group of trajectory lines. In this case, the first marker does not exist, and the specified marker obtained by the electronic device based on the n trajectory lines only includes the second marker.
[0193] (2) The electronic device performs polygon fitting processing on the first set of trajectory lines. If a polygon is fitted through at least one trajectory line in the first set of trajectory lines, the electronic device determines the fitted polygon as the first marker. If multiple polygons are fitted through at least one trajectory line in the first set of trajectory lines, the electronic device determines the first marker based on the overlap between the fitted multiple polygons. If no polygon is fitted through any of the trajectory lines in the first set of trajectory lines, the electronic device determines that there is no first marker.
[0194] In this embodiment of the application, performing polygon fitting processing on the first group of trajectory lines is to attempt to fit one or more polygons using one or more trajectory lines from the first group of trajectory lines.
[0195] If a polygon is fitted by at least one of the trajectory lines in the first set, it means that the at least one trajectory line forms a closed figure (i.e., the polygon), and therefore the polygon formed can be identified as the first marker.
[0196] The operation of the electronic device to determine the first mark based on the fitted polygons and their overlap can be as follows: if none of the polygons overlap, then the polygons are determined as the first mark. If there are overlapping polygons among the polygons, then the overlapping polygons are merged to obtain a new polygon, and the merged new polygon and the polygons among the polygons that do not overlap with other polygons are determined as the first mark.
[0197] If multiple polygons are fitted using at least one trajectory line from the first set of trajectory lines, it means that the at least one trajectory line forms multiple closed figures (i.e., the multiple polygons). Therefore, the first marker can be determined based on the overlap between these multiple polygons. Specifically, for any two polygons among these multiple polygons, since both polygons are used to indicate the content to be identified, if these two polygons overlap, they can be merged to obtain a new polygon as the first marker. The area enclosed by the new polygon is the union of the areas enclosed by the two polygons.
[0198] If no polygon is fitted by the first set of trajectory lines, it means that the first set of trajectory lines forms non-closed figures, and therefore it can be determined that there is no first marker.
[0199] (3) The electronic device performs polygon fitting processing on the second set of trajectory lines. If a polygon is fitted through at least one trajectory line in the second set of trajectory lines, the electronic device determines the fitted polygon as the second marker. If multiple polygons are fitted through at least one trajectory line in the second set of trajectory lines, the electronic device determines the second marker based on the overlap between the fitted multiple polygons. If no polygon is fitted through any of the trajectory lines in the second set of trajectory lines, the electronic device determines that no second marker exists.
[0200] In this embodiment of the application, performing polygon fitting processing on the second set of trajectory lines is to attempt to fit one or more polygons using one or more trajectory lines from the second set of trajectory lines.
[0201] If a polygon is fitted by at least one of the trajectory lines in the second set, it means that the at least one trajectory line forms a closed figure (i.e., the polygon), and therefore the polygon formed can be identified as the second marker.
[0202] The operation of the electronic device to determine the second mark based on the fitted polygons and their overlap can be as follows: if none of the polygons overlap, then the polygons are determined as the second mark. If there are overlapping polygons among the polygons, then the overlapping polygons are merged to obtain a new polygon, and the merged new polygon and the polygons among the polygons that do not overlap with other polygons are determined as the second mark.
[0203] If multiple polygons are fitted using at least one of the trajectory lines in the second set, it means that the at least one trajectory line forms multiple closed figures (i.e., the multiple polygons). Therefore, the second mark can be determined based on the overlap between these multiple polygons. Specifically, for any two polygons among these multiple polygons, since both polygons are used to indicate content that does not need to be identified, if these two polygons overlap, they can be merged to obtain a new polygon as the second mark. The area enclosed by the new polygon is the union of the areas enclosed by the two polygons.
[0204] If no polygon is fitted by the second set of trajectory lines, it means that the second set of trajectory lines forms non-closed figures, and therefore it can be determined that there is no second marker.
[0205] In the fourth possible implementation, if the difference between the first mark and the second mark is that the first mark is a closed shape and the second mark is also a closed shape, but the line colors of the first mark and the second mark are different, then the operation of step 1202 may include the following steps (1) to (3):
[0206] (1) The electronic device groups the n trajectory lines to obtain the first group of trajectory lines and / or the second group of trajectory lines.
[0207] The line color of the first group of trajectory lines is the same as the line color of the first marker, and the line color of the second group of trajectory lines is the same as the line color of the second marker.
[0208] Specifically, if the line color of each of the n trajectory lines is the same as the line color of the first marker, then after the electronic device groups the n trajectory lines, it will only obtain the first group of trajectory lines. In this case, there is no second marker, and the specified marker obtained by the electronic device based on the n trajectory lines only includes the first marker.
[0209] If the line color of each of the n trajectory lines is the same as the line color of the second marker, then after the electronic device groups the n trajectory lines, it will only obtain the second group of trajectory lines. In this case, the first marker does not exist, and the specified marker obtained by the electronic device based on the n trajectory lines only includes the second marker.
[0210] (2) The electronic device performs polygon fitting processing on the first set of trajectory lines. If a polygon is fitted through at least one trajectory line in the first set of trajectory lines, the electronic device determines the fitted polygon as the first marker. If multiple polygons are fitted through at least one trajectory line in the first set of trajectory lines, the electronic device determines the first marker based on the overlap between the fitted multiple polygons. If no polygon is fitted through any of the trajectory lines in the first set of trajectory lines, the electronic device determines that there is no first marker.
[0211] (3) The electronic device performs polygon fitting processing on the second set of trajectory lines. If a polygon is fitted through at least one trajectory line in the second set of trajectory lines, the electronic device determines the fitted polygon as the second marker. If multiple polygons are fitted through at least one trajectory line in the second set of trajectory lines, the electronic device determines the second marker based on the overlap between the fitted multiple polygons. If no polygon is fitted through any of the trajectory lines in the second set of trajectory lines, the electronic device determines that no second marker exists.
[0212] Step 1203: The electronic device identifies the target content from the screen content of the electronic device based on the specified marker.
[0213] In this embodiment, it is not necessary to recognize the entire screen content of the electronic device. Instead, target content that meets the user's needs can be identified from the screen content based on the user's specified markers. The recognition time is short, content recognition can be completed quickly, and power consumption can be saved. Furthermore, since the user can select which content to recognize or not recognize according to their needs—that is, the user can select a local area of content on the screen for recognition or not, or select multiple discontinuous areas of content on the screen for recognition or not—this embodiment can achieve cross-line recognition of screen content, making screen content recognition more flexible.
[0214] It should be noted that the specified markers acquired by the electronic device may include one or more first markers, or one or more second markers. That is, the user can draw one or more first markers on the screen to indicate one or more pieces of content that need to be recognized. The user can also draw one or more second markers on the screen to indicate one or more pieces of content that do not need to be recognized.
[0215] Optionally, the first marker is used to delineate the content to be identified, and in this case, the first marker is a closed shape. The second marker is used to delineate the content that does not need to be identified, and in this case, the second marker is a closed shape; or, the second marker is used to obscure the content that does not need to be identified, and in this case, the second marker is a non-closed shape.
[0216] In this case, step 1203 can include the following three methods:
[0217] In the first method, if the specified mark includes the first mark but does not include the second mark, the electronic device identifies the target content from the screen content of the electronic device based on the area enclosed by the first mark (which may be referred to as the first area).
[0218] The operation of an electronic device identifying target content from the screen content of an electronic device based on a first region defined by a first marker can include the following four cases:
[0219] In the first scenario: If the content within the first area of the electronic device's screen is text, the electronic device determines the text paragraph containing that text as the first paragraph. Based on the coordinates of each sentence in the first paragraph and the coordinate range of the first area, the electronic device determines the overlap between each sentence in the first paragraph and the first area. Based on this overlap, the electronic device identifies the target sentence from the first paragraph; the target sentence is the target content. Alternatively, the electronic device determines the overlap between each character in the first paragraph and the first area based on the coordinates of each character in the first paragraph and the coordinate range of the first area. Based on this overlap, the electronic device identifies the target character from the first paragraph; the target character is the target content.
[0220] The electronic device first determines the coordinate range of the first region defined by the first marker. Then, it detects the content on the screen that falls within the coordinate range of the first region to determine whether the content in the first region is text, a table, an image, or a formula. If the content in the first region is text, the electronic device performs line-by-line detection on the screen to determine the coordinates of each line (i.e., the coordinates of the start and end positions of the text lines). Then, based on the coordinates of each line, it performs text segment analysis to obtain the coordinates of each text segment (i.e., the coordinates of the start and end positions of the text segments) and the coordinates of each line within each segment. Afterward, the electronic device can determine the text segment (i.e., the first segment) containing the text content in the first region based on the coordinate range of the first region, the coordinates of each text segment, and the coordinates of each line within each segment. At this point, the coordinates of the first segment and the coordinates of each line within the first segment are also determined.
[0221] As an example, the electronic device can determine the coordinates of each sentence in the first paragraph. For instance, the electronic device can use a sentence segmentation algorithm in natural language processing (NLP) to determine the coordinates of each sentence in the first paragraph. Then, the electronic device can determine the overlap between each sentence in the first paragraph and the first region based on the coordinates of each sentence in the first paragraph and the coordinate range of the first region. Optionally, the overlap between a sentence in the first paragraph and the first region can be the ratio of the overlapping area of the sentence (i.e., the area of overlap between the sentence and the first region) to the total area of the sentence. In other words, the overlap between the sentence and the first region represents what proportion of the sentence's area falls within the first region. For any sentence in the first paragraph, if the overlap between the sentence and the first region is greater than or equal to a first overlap, the electronic device can identify the sentence as the target sentence and perform recognition on it. The first overlap can be preset, and the first overlap can be set relatively large, such as 50% or 60%, etc. This embodiment does not limit this.
[0222] As another example, the electronic device can determine the coordinates of each character in the first paragraph. These characters may include letters, numbers, Chinese characters, symbols, etc. The electronic device can then determine the overlap between each character in the first paragraph and the first region based on the coordinates of the characters in the first paragraph and the coordinate range of the first region. Optionally, the overlap between a character in the first paragraph and the first region can be the ratio of the overlapping area of the character (i.e., the area of overlap between the character and the first region) to the total area of the character. In other words, the overlap between the character and the first region represents the proportion of the character's area that falls within the first region. For any character in the first paragraph, if the overlap between the character and the first region is greater than or equal to a first overlap, the electronic device can identify this character as the target character and perform recognition on it.
[0223] For example, such as Figure 17As shown, the electronic device can determine the coordinates of each character in the text paragraph (i.e., the first paragraph) where the text content in the first area is located. Then, the electronic device can determine the overlap degree of each character in the first paragraph with the first area according to the coordinates of each character in the first paragraph and the coordinate range of the first area. In this case, since the overlap degree of other characters in the first paragraph except the three characters "大", "之", and "组" with the first area is 100%, the electronic device can determine other characters in the first paragraph except the three characters "大", "之", and "组" as target characters and identify these characters. Since the overlap degree of the character "组" in the first paragraph with the first area is greater than 50%, the electronic device can determine the character "组" in the first paragraph as a target character and identify this character. However, since the overlap degrees of the two characters "大" and "之" in the first paragraph with the first area are both less than 50%, the electronic device does not identify the two characters "大" and "之" in the first paragraph.
[0224] It should be noted that in the embodiments of the present application, the actual勾画 area (i.e., the first area) can be combined with the text paragraph, and then the overlap situation between the sentences or characters in the text paragraph and the勾画 area can be analyzed by using the sentence coordinates or character coordinates of the text paragraph, so as to determine the target sentences or target characters to be recognized accordingly. In this way, even when the user勾画s an incomplete sentence or勾画s half a character, the user's勾画 intention can be restored to a certain extent, so that the target sentences or target characters that meet the user's needs can be accurately recognized.
[0225] Further, after the electronic device recognizes the target sentence from the first paragraph, it also obtains the encoding information of the target sentence. In this case, the electronic device can also extract the recognized target sentence. Optionally, the electronic device can store the recognized target sentence (i.e., the encoding information of the target sentence) in the system clipboard.
[0226] Similarly, after the electronic device recognizes the target character from the first paragraph, it also obtains the encoding information of the target character. In this case, the electronic device can also extract the recognized target character. Optionally, the electronic device can store the recognized target character (i.e., the encoding information of the target character) in the system clipboard.
[0227] It should be noted that since the embodiments of the present application can achieve跨行 recognition of the screen content of the electronic device, after extracting the recognized target content,跨行 extraction of the screen content of the electronic device is achieved, thereby improving the flexibility of screen content extraction.
[0228] The system clipboard is a data storage area used to temporarily store data that may need to be pasted later.
[0229] As an example, if an edit box exists on the screen of an electronic device, the device can directly paste the target content (target sentence or target character) stored in the system clipboard into that edit box.
[0230] As another example, if an edit box exists on the screen of an electronic device and that edit box has focus, the electronic device can paste the target content stored in the system clipboard into that edit box.
[0231] It is worth noting that the edit box that gains focus is most likely the edit box that the user currently needs to use. Therefore, in this embodiment, the identified target content can be directly pasted into the edit box that gains focus, which makes it more convenient for the user and improves the user experience.
[0232] As another example, once an electronic device detects a paste command in any edit box, it can paste the target content stored in the system clipboard into that edit box.
[0233] As another example, once an electronic device detects that any edit box has focus, it can paste the target content stored in the system clipboard into that edit box.
[0234] For example, such as Figure 18 As shown in Figure (a), after the user marks a first mark on the screen content 1801 of the electronic device, the electronic device stores the target content in the first area defined by the first mark into the system clipboard. Then, as... Figure 18 As shown in Figure (b), the user switches from screen content 1801 of the electronic device to another interface 1802, which contains an edit box. When the user clicks on the edit box in interface 1802, the edit box gains focus. In this case, as shown in Figure (b), the user can switch from screen content 1801 to another interface 1802. Figure 18 As shown in Figure (c), the electronic device will automatically paste the target content stored in the system clipboard into the edit box.
[0235] The second scenario: If the content on the screen of the electronic device located in the first area is table content, then the electronic device determines that the table content in the first area belongs to the first table. The electronic device determines the overlap between each cell in the first table and the first area based on the coordinates of each cell in the first table and the coordinate range of the first area; based on the overlap between each cell in the first table and the first area, it identifies the text in the target cell from the first table, and the text in the target cell is the target content.
[0236] The electronic device first determines the coordinate range of the first region defined by the first marker. Then, it detects the content on the screen of the electronic device that falls within the coordinate range of the first region to determine whether the content in the first region is text, table content, image content, or formula content. If the content in the first region is table content, the electronic device analyzes the table structure of the table to which the table content in the first region belongs (i.e., the first table) to obtain the coordinates of each cell in the first table. Then, the electronic device can determine the overlap between each cell in the first table and the first region based on the coordinate range of the first region. Optionally, the overlap between a cell in the first table and the first region can be the ratio of the overlapping area of the cell (i.e., the area of overlap between the cell and the first region) to the total area of the cell. In other words, the overlap between the cell and the first region represents what proportion of the cell's area falls within the first region. For any cell in the first table, if the overlap between the cell and the first region is greater than or equal to a second overlap, the electronic device can identify this cell as the target cell and recognize the text in this cell. The second overlap can be preset, for example, the second overlap can be 30%, 40%, etc., and this application embodiment does not limit it.
[0237] In some embodiments, after the electronic device determines the target cell based on the overlap between each cell in the first table and the first region, it can also perform coordinate alignment on all the determined target cells in the first table. That is, it can perform vertical and horizontal alignment on all the determined target cells in the first table to obtain a sub-table containing all the determined target cells in rows a and columns b from the first table, where a and b are both positive integers.
[0238] Then, the electronic device can recognize the text in the target cell of the sub-table. In this case, the target cell in the sub-table will contain text, while other cells will be empty. Alternatively, the electronic device can identify all cells in the sub-table as target cells and then recognize the text in all cells of the sub-table. In this case, all cells in the sub-table will contain text.
[0239] For example, such as Figure 19As shown in Figure (a), the electronic device can determine the coordinates of each cell in the table to which the table content in the first region belongs (i.e., the first table). Then, the electronic device can determine the overlap between each cell in the first table and the first region based on the coordinates of each cell in the first table and the coordinate range of the first region. In this case, since the overlap between the cell in row 2, column 2, row 2, column 3, and row 3, column 2 of the first table and the first region is greater than 30%, the electronic device identifies these three cells as target cells. Then, the electronic device aligns the coordinates of these three target cells in the first table to obtain a 3-row, 2-column sub-table from the first table. As an example, the electronic device recognizes the text in the target cells of this sub-table, and the recognized sub-table can be as follows: Figure 19 As shown in Figure (b), in this case, the target cell in the identified sub-table contains text, while the other cells are empty. As another example, the electronic device identifies all cells in the sub-table as target cells and recognizes the text in all cells of the sub-table; the identified sub-table can be as follows: Figure 19 As shown in Figure (c), in this case, all cells in the identified sub-table contain text.
[0240] In this embodiment, the actual drawn area (i.e., the first area) can be combined with a table. Then, the overlap between the table cells and the drawn area can be analyzed using the table's cell coordinates to determine the target cell to be identified. This way, even if the user doesn't draw the entire cell, the user's drawing intention can be reconstructed to a certain extent, thus accurately identifying the text in the target cell that meets the user's needs.
[0241] Furthermore, after the electronic device identifies the sub-table, it obtains the sub-table's encoding information. In this case, the electronic device can also extract the identified sub-table. Optionally, the electronic device can store the identified sub-table (i.e., the sub-table's encoding information) to the system clipboard.
[0242] As an example, if an electronic device's screen contains an edit box, the device can directly paste a sub-table stored in the system clipboard into that edit box.
[0243] As another example, if an edit box exists on the screen of an electronic device and that edit box has focus, the electronic device can paste a sub-table stored in the system clipboard into that edit box.
[0244] As another example, when an electronic device detects a paste command in any edit box, it can paste a sub-table stored in the system clipboard into that edit box.
[0245] As another example, once an electronic device detects that any edit box has focus, it can paste a sub-table stored in the system clipboard into that edit box.
[0246] The third scenario: If the content in the first area of the electronic device's screen is an image, the electronic device determines that the image content in the first area belongs to the first image, identifies the first image based on its coordinates, and identifies the first image as the target content.
[0247] The electronic device can first determine the coordinate range of the first region defined by the first marker, and then detect the content on the screen of the electronic device that falls within the coordinate range of the first region to determine whether the content in the first region is text, table content, image content, or formula content. If the content in the first region is image content, the electronic device can analyze the image features of the image to which the image content in the first region belongs (i.e., the first image) to obtain the coordinates of the first image. Then, the electronic device can identify the first image based on the coordinates of the first image.
[0248] When the electronic device identifies the first image based on its coordinates, it can also correct the identified first image to adjust it to normal if the first image is skewed or distorted.
[0249] For example, such as Figure 20 As shown in Figure (a), the electronic device can determine the coordinates of the image to which the image content in the first region belongs (i.e., the first image), then identify the first image based on its coordinates, and correct the identified first image to restore it to its original shape. Figure 20 The normal image shown in Figure (b) is shown in the image.
[0250] Furthermore, after the electronic device recognizes the first image, it obtains the encoded information of the first image. In this case, the electronic device can also extract the recognized first image. Optionally, the electronic device can store the recognized first image (i.e., the encoded information of the first image) to the system clipboard.
[0251] As an example, if an edit box exists on the screen of an electronic device, the device can directly paste the first image stored in the system clipboard into that edit box.
[0252] As another example, if an edit box exists on the screen of an electronic device and that edit box is focused, the electronic device can paste the first image stored in the system clipboard into that edit box.
[0253] As another example, once an electronic device detects a paste command in any edit box, it can paste the first image stored in the system clipboard into that edit box.
[0254] As another example, once an electronic device detects that any edit box has focus, it can paste the first image stored in the system clipboard into that edit box.
[0255] The fourth scenario: If the content in the first area of the screen of the electronic device is formula content, then the electronic device identifies the formula (which can be called the first formula) in the first area, and the first formula is the target content.
[0256] The electronic device can first determine the coordinate range of the first region defined by the first marker, and then detect the content on the screen of the electronic device that falls within the coordinate range of the first region to determine whether the content in the first region is text, table content, image content, or formula content. If the content in the first region is formula content, the electronic device can directly identify the formula (i.e., the first formula) contained in the formula content in the first region.
[0257] For example, such as Figure 21 As shown in Figure (a), the electronic device can directly identify the formula (i.e., the first formula) to which the formula content in the first area belongs. The identified first formula can be... Figure 21 The formula is shown in Figure (b).
[0258] Furthermore, after the electronic device identifies the first formula, it obtains the encoding information of the first formula. In this case, the electronic device can also extract the identified first formula. Optionally, the electronic device can store the identified first formula (i.e., the encoding information of the first formula) to the system clipboard.
[0259] As an example, if an edit box exists on the screen of an electronic device, the device can directly paste the first formula stored in the system clipboard into that edit box.
[0260] As another example, if an edit box exists on the screen of an electronic device and that edit box is focused, the electronic device can paste the first formula stored in the system clipboard into that edit box.
[0261] As another example, once an electronic device detects a paste command in any edit box, it can paste the first formula stored in the system clipboard into that edit box.
[0262] As another example, once the electronic device detects that any edit box has focus, it can paste the first formula stored in the system clipboard into that edit box.
[0263] In the second method, if the specified mark includes the second mark but does not include the first mark, the electronic device identifies the target content from the screen content of the electronic device based on the area enclosed by the second mark (which may be referred to as the second region) or based on the content obscured by the second mark.
[0264] If the second mark is used to delineate content that does not need to be recognized, the electronic device recognizes the target content from the screen content of the electronic device based on the area delineated by the second mark (i.e., the second region). If the second mark is used to obscure content that does not need to be recognized, the electronic device recognizes the target content from the screen content of the electronic device based on the content obscured by the second mark.
[0265] The operation of an electronic device identifying target content from the screen content of an electronic device based on a second region defined by a second mark can include the following four cases:
[0266] Scenario 1: If the content located in the second area of the electronic device's screen is text content, the electronic device determines the text paragraph containing all other text content besides the text content in the second area as the second paragraph. The electronic device determines the overlap between each sentence in the second paragraph and the second area based on the coordinates of each sentence in the second paragraph and the coordinate range of the second area. Based on this overlap, the electronic device identifies the target sentence from the second paragraph; the target sentence is the target content. Alternatively, the electronic device determines the overlap between each character in the second paragraph and the second area based on the coordinates of each character in the second paragraph and the coordinate range of the second area. Based on this overlap, the electronic device identifies the target character from the second paragraph; the target character is the target content.
[0267] The electronic device first determines the coordinate range of the second region defined by the second marker. Then, it detects the content on the screen that falls within the coordinate range of the second region to determine whether the content in the second region is text, a table, an image, or a formula. If the content in the second region is text, the electronic device performs line-by-line detection on the screen to determine the coordinates of each line. Then, based on the coordinates of each line, it performs text segment analysis to obtain the coordinates of each text segment and the coordinates of each line within each segment. Afterward, based on the coordinate range of the second region, the coordinates of each text segment, and the coordinates of each line within each segment, the electronic device can determine the text segment containing all other text content besides the text in the second region (i.e., the second segment). At this point, the coordinates of the second segment and the coordinates of each line within the second segment are also determined.
[0268] As an example, the electronic device can determine the coordinates of each sentence in the second paragraph. For instance, the electronic device can use a sentence segmentation algorithm in NLP to determine the coordinates of each sentence in the second paragraph. Then, the electronic device can determine the overlap between each sentence in the second paragraph and the second region based on the coordinates of each sentence in the second paragraph and the coordinate range of the second region. Optionally, the overlap between a sentence in the second paragraph and the second region can be the ratio of the overlapping area of the sentence (i.e., the area of overlap between the sentence and the second region) to the total area of the sentence. In other words, the overlap between the sentence and the second region represents what proportion of the sentence's area falls within the second region. For any sentence in the second paragraph, if the overlap between the sentence and the second region is less than a first overlap, the electronic device can identify the sentence as the target sentence and perform recognition on it. The first overlap can be preset, and the first overlap can be set relatively large, such as 50% or 60%, etc. This application embodiment does not limit this.
[0269] As another example, the electronic device can determine the coordinates of each character in the second paragraph. These characters can include letters, numbers, Chinese characters, symbols, etc. The electronic device can then determine the overlap between each character in the second paragraph and the second region based on the coordinates of the characters in the second paragraph and the coordinate range of the second region. Optionally, the overlap between a character in the second paragraph and the second region can be the ratio of the overlapping area of that character (i.e., the area of overlap between that character and the second region) to the total area of that character. In other words, the overlap between that character and the second region represents what proportion of the character's area falls within the second region. For any character in the second paragraph, if the overlap between that character and the second region is less than a first overlap, the electronic device can identify that character as the target character and perform recognition on it.
[0270] Furthermore, after identifying the target sentence from the second paragraph, the electronic device obtains the encoded information of the target sentence. In this case, the electronic device can also extract the identified target sentence. Optionally, the electronic device can store the identified target sentence (i.e., the encoded information of the target sentence) to the system clipboard.
[0271] Similarly, after the electronic device identifies the target character from the second paragraph, it obtains the encoding information of the target character. In this case, the electronic device can also extract the identified target character. Optionally, the electronic device can store the identified target character (i.e., the encoding information of the target character) to the system clipboard.
[0272] As an example, if an edit box exists on the screen of an electronic device, the device can directly paste the target content (target sentence or target character) stored in the system clipboard into that edit box.
[0273] As another example, if an edit box exists on the screen of an electronic device and that edit box has focus, the electronic device can paste the target content stored in the system clipboard into that edit box.
[0274] As another example, once an electronic device detects a paste command in any edit box, it can paste the target content stored in the system clipboard into that edit box.
[0275] As another example, once an electronic device detects that any edit box has focus, it can paste the target content stored in the system clipboard into that edit box.
[0276] The second scenario: If the content on the screen of the electronic device located in the second area is table content, then the electronic device determines that the table content in the second area belongs to the second table. The electronic device determines the overlap between each cell in the second table and the second area based on the coordinates of each cell in the second table and the coordinate range of the second area; based on the overlap between each cell in the second table and the second area, it identifies the text in the target cell from the second table, and the text in the target cell is the target content.
[0277] The electronic device first determines the coordinate range of the second region defined by the second marker. Then, it detects the content on the screen that falls within the coordinate range of the second region to determine whether the content in the second region is text, table content, image content, or formula content. If the content in the second region is table content, the electronic device analyzes the table structure of the table to which the table content in the second region belongs (i.e., the second table) to obtain the coordinates of each cell in the second table. Then, the electronic device can determine the overlap between each cell in the second table and the second region based on the coordinate range of the second region. Optionally, the overlap between a cell in the second table and the second region can be the ratio of the overlapping area of the cell (i.e., the area of overlap between the cell and the second region) to the total area of the cell. In other words, the overlap between the cell and the second region represents what percentage of the cell's area falls within the second region. For any cell in the second table, if the overlap between the cell and the second region is less than the second overlap, the electronic device can identify this cell as the target cell and recognize the text in this cell. The second overlap can be preset, for example, the second overlap can be 30%, 40%, etc., and this application embodiment does not limit it.
[0278] In some embodiments, after the electronic device determines the target cells based on the overlap between each cell in the second table and the second region, it can further align all the determined target cells in the second table, that is, align all the determined target cells vertically and horizontally in the second table to obtain a sub-table containing all the determined target cells in rows a and columns b, where a and b are both positive integers. Then, the electronic device can recognize the text in the target cells of this sub-table. In this case, the recognized target cells in the sub-table contain text, while other cells are empty.
[0279] For example, such as Figure 22 As shown in Figure (a), the electronic device can determine the coordinates of each cell in the table to which the table content in the second area belongs (i.e., the second table). Then, the electronic device can determine the overlap between each cell in the second table and the second area based on the coordinates of each cell in the second table and the coordinate range of the second area. In this case, since the overlap between the cells in the second table other than the cells in the 2nd row, 2nd column, 2nd row, 3rd column, and 3rd row, 2nd column and the second area is less than 30%, the electronic device identifies all other cells in the second table except these three as target cells. Next, the electronic device aligns the coordinates of all target cells in the second table to obtain a 5-row, 3-column sub-table. The electronic device then recognizes the text in the target cells of this sub-table, and the recognized sub-table can be displayed as follows: Figure 22 As shown in Figure (b), in this case, the target cell in the identified sub-table contains text, while the other cells are empty.
[0280] Furthermore, after the electronic device identifies the sub-table, it obtains the sub-table's encoding information. In this case, the electronic device can also extract the identified sub-table. Optionally, the electronic device can store the identified sub-table (i.e., the sub-table's encoding information) to the system clipboard.
[0281] As an example, if an electronic device's screen contains an edit box, the device can directly paste a sub-table stored in the system clipboard into that edit box.
[0282] As another example, if an edit box exists on the screen of an electronic device and that edit box has focus, the electronic device can paste a sub-table stored in the system clipboard into that edit box.
[0283] As another example, when an electronic device detects a paste command in any edit box, it can paste a sub-table stored in the system clipboard into that edit box.
[0284] As another example, once an electronic device detects that any edit box has focus, it can paste a sub-table stored in the system clipboard into that edit box.
[0285] The third scenario: If the content located in the second area of the electronic device's screen content is image content, then the electronic device determines that the image content in the second area belongs to the second image. Based on the coordinates of the second image, it identifies other content from the screen content of the electronic device besides the second image. The other content from the screen content of the electronic device besides the second image is the target content.
[0286] The electronic device can first determine the coordinate range of the second region defined by the second marker, and then detect the content on the screen that falls within the coordinate range of the second region to determine whether the content in the second region is text, table content, image content, or formula content. If the content in the second region is an image, the electronic device can analyze the image features of the image to which the image content in the second region belongs (i.e., the second image) to obtain the coordinates of the second image. Then, the electronic device can identify other content from the screen content other than the second image based on the coordinates of the second image.
[0287] The fourth scenario: If the content located in the second area of the screen content of the electronic device is formula content, then the electronic device identifies other content in the screen content of the electronic device other than the formula (which can be called the second formula) located in the formula content of the second area, and the other content in the screen content of the electronic device other than the second formula is the target content.
[0288] The electronic device can first determine the coordinate range of the second region defined by the second marker, and then detect the content on the screen of the electronic device that falls within the coordinate range of the second region to determine whether the content in the second region is text, table content, image content, or formula content. If the content in the second region is formula content, the electronic device can directly identify other content on the screen of the electronic device besides the formula content in the second region.
[0289] The operation of an electronic device identifying target content from its screen content based on the content obscured by the second marker can include the following four scenarios:
[0290] In the first scenario: if the content on the screen of the electronic device that is obscured by the second marker is text content, then the electronic device determines the text paragraph containing all other text content on the screen except for the text content obscured by the second marker as the second paragraph. The electronic device determines the degree of occlusion of each sentence in the second paragraph based on the coordinates of each sentence and the coordinates of the second marker, and identifies the target sentence from the second paragraph based on the degree of occlusion; the target sentence is the target content. Alternatively, the electronic device determines the degree of occlusion of each character in the second paragraph based on the coordinates of each character and the coordinates of the second marker, and identifies the target character from the second paragraph based on the degree of occlusion; the target character is the target content.
[0291] The electronic device first determines the coordinates of the second marker, then detects the content on the screen located at the coordinates of the second marker to determine whether the content obscured by the second marker is text, table content, image content, or formula content. If the content obscured by the second marker is text, the electronic device performs text line detection on the screen to determine the coordinates of each text line. Then, based on the coordinates of each text line, it performs text segment analysis to obtain the coordinates of each text segment and the coordinates of each text line within each text segment. Afterward, based on the coordinates of the second marker, the coordinates of each text segment, and the coordinates of each text line within each text segment, the electronic device can determine the text segment containing all text content except the text obscured by the second marker (i.e., the second segment). At this point, the coordinates of the second segment and the coordinates of each text line within the second segment are also determined.
[0292] As an example, the electronic device can determine the coordinates of each sentence in the second paragraph. For instance, the electronic device can use a sentence segmentation algorithm in NLP to determine the coordinates of each sentence in the second paragraph. Then, the electronic device can determine the occlusion degree of each sentence in the second paragraph based on the coordinates of each sentence and the coordinates of the second marker. Optionally, the occlusion degree of a sentence in the second paragraph can be the ratio of the occlusion area of the sentence (i.e., the area of the sentence occluded by the second marker) to the total area of the sentence. In other words, the occlusion degree of the sentence represents what proportion of the sentence's area is occluded by the second marker. For any sentence in the second paragraph, if the occlusion degree of the sentence is less than the first occlusion degree, the electronic device can identify the sentence as the target sentence and perform recognition on it. The first occlusion degree can be preset, for example, it can be 30%, 40%, etc., and this embodiment does not limit this.
[0293] As another example, the electronic device can determine the coordinates of each character in the second paragraph. These characters may include letters, numbers, Chinese characters, symbols, etc. The electronic device can then determine the occlusion degree of each character in the second paragraph based on the coordinates of the characters and the coordinates of the second marker. Optionally, the occlusion degree of a character in the second paragraph can be the ratio of the occluded area of that character (i.e., the area of that character occluded by the second marker) to the total area of that character. In other words, the occlusion degree represents what proportion of the character's area is occluded by the second marker. For any character in the second paragraph, if the occlusion degree of that character is less than the first occlusion degree, the electronic device can identify that character as the target character and perform recognition on it.
[0294] Furthermore, after identifying the target sentence from the second paragraph, the electronic device obtains the encoded information of the target sentence. In this case, the electronic device can also extract the identified target sentence. Optionally, the electronic device can store the identified target sentence (i.e., the encoded information of the target sentence) to the system clipboard.
[0295] Similarly, after the electronic device identifies the target character from the second paragraph, it obtains the encoding information of the target character. In this case, the electronic device can also extract the identified target character. Optionally, the electronic device can store the identified target character (i.e., the encoding information of the target character) to the system clipboard.
[0296] As an example, if an edit box exists on the screen of an electronic device, the device can directly paste the target content (target sentence or target character) stored in the system clipboard into that edit box.
[0297] As another example, if an edit box exists on the screen of an electronic device and that edit box has focus, the electronic device can paste the target content stored in the system clipboard into that edit box.
[0298] As another example, once an electronic device detects a paste command in any edit box, it can paste the target content stored in the system clipboard into that edit box.
[0299] As another example, once an electronic device detects that any edit box has focus, it can paste the target content stored in the system clipboard into that edit box.
[0300] The second scenario: If the content on the screen of the electronic device that is obscured by the second marker is table content, then the electronic device determines that the table content obscured by the second marker belongs to the second table. The electronic device determines the degree of occlusion of each cell in the second table based on the coordinates of each cell and the coordinates of the second marker; based on the degree of occlusion of each cell in the second table, it identifies the text in the target cell, and the text in the target cell is the target content.
[0301] The electronic device first determines the coordinates of the second marker, then detects the content on the screen located at those coordinates to determine whether the content obscured by the second marker is text, table content, image content, or formula content. If the obscured content is table content, the electronic device analyzes the table structure of the table (i.e., the second table) to obtain the coordinates of each cell in the second table. The electronic device then determines the degree of occlusion of each cell in the second table based on the coordinates of the cells and the second marker. Optionally, the degree of occlusion of a cell in the second table can be the ratio of the area obscured by the second marker to the total area of the cell; that is, the degree of occlusion represents the proportion of the cell's area obscured by the second marker. For any cell in the second table, if the degree of occlusion is less than the second occlusion degree, the electronic device can identify this cell as the target cell and recognize the text within it. The second overlap can be preset, for example, the second overlap can be 10%, 20%, etc., and this application embodiment does not limit it.
[0302] In some embodiments, after the electronic device determines the target cell based on the occlusion degree of each cell in the second table, it can further align all the determined target cells in the second table, that is, align all the determined target cells vertically and horizontally in the second table to obtain a sub-table containing all the determined target cells in row a and column b, where a and b are both positive integers. Then, the electronic device can recognize the text in the target cells of this sub-table. In this case, the recognized target cells in the sub-table contain text, while other cells are empty.
[0303] Furthermore, after the electronic device identifies the sub-table, it obtains the sub-table's encoding information. In this case, the electronic device can also extract the identified sub-table. Optionally, the electronic device can store the identified sub-table (i.e., the sub-table's encoding information) to the system clipboard.
[0304] As an example, if an electronic device's screen contains an edit box, the device can directly paste a sub-table stored in the system clipboard into that edit box.
[0305] As another example, if an edit box exists on the screen of an electronic device and that edit box has focus, the electronic device can paste a sub-table stored in the system clipboard into that edit box.
[0306] As another example, when an electronic device detects a paste command in any edit box, it can paste a sub-table stored in the system clipboard into that edit box.
[0307] As another example, once an electronic device detects that any edit box has focus, it can paste a sub-table stored in the system clipboard into that edit box.
[0308] The third scenario: If the content on the screen of the electronic device that is obscured by the second mark is image content, then the electronic device determines that the image content obscured by the second mark belongs to the second image. Based on the coordinates of the second image, the electronic device identifies other content on the screen of the electronic device that is not the second image. The other content on the screen of the electronic device that is not the second image is the target content.
[0309] The electronic device can first determine the coordinates of the second marker, and then detect the content on the screen located at the coordinates of the second marker to determine whether the content obscured by the second marker is text, table content, image content, or formula content. If the content obscured by the second marker is image content, the electronic device can analyze the image features of the image to which the obscured image content belongs (i.e., the second image) to obtain the coordinates of the second image. Then, the electronic device can identify other content besides the second image from the screen content based on the coordinates of the second image.
[0310] The fourth scenario: If the content on the screen of the electronic device that is obscured by the second mark is formula content, then the electronic device identifies the other content on the screen of the electronic device that is not the formula content obscured by the second mark (which can be called the second formula), and the other content on the screen of the electronic device that is not the second formula is the target content.
[0311] The electronic device can first determine the coordinates of the second marker, and then detect the content on the screen located at the coordinates of the second marker to determine whether the content obscured by the second marker is text, table content, image content, or formula content. If the content obscured by the second marker is formula content, the electronic device can directly identify other content on the screen besides the formula containing the obscured formula content.
[0312] The third method: If the specified marker includes the first marker and the second marker, the electronic device can identify the target content from the screen content of the electronic device through the following methods 1 to 6.
[0313] Method 1: If the first and second markers do not overlap, the electronic device can retain the first marker and discard the second marker in the specified markers, and then identify the target content from the screen content of the electronic device using the first method described above. Alternatively, if the first and second markers do not overlap, the electronic device can discard the first marker and retain the second marker in the specified markers, and then identify the target content from the screen content of the electronic device using the second method described above.
[0314] In this embodiment, if both the first marker and the second marker exist but do not overlap, their delineation intentions conflict. In this case, one can be discarded while the other is retained so that content recognition can continue normally.
[0315] For example, such as Figure 23 As shown, the electronic device acquires a specified tag, assuming the specified tag is as follows: Figure 24 Figure (a) shows a first mark and a second mark. The electronic device then determines the overlap between the first mark and the second mark, because... Figure 24 In Figure (a), the first and second marks do not overlap, so the electronic device can determine that the overlap between the first and second marks is: the first and second marks do not overlap. In this case, the electronic device can retain the first mark and discard the second mark, obtaining... Figure 24 The specified marker shown in Figure (b) includes the first marker but excludes the second marker, thus allowing the target content to be identified from the screen content of the electronic device using the first method described above. Alternatively, the electronic device can discard the first marker and retain the second marker, obtaining... Figure 24 The specified marker shown in Figure (c) includes the second marker but does not include the first marker, so the target content can be identified from the screen content of the electronic device in the second way described above.
[0316] Method 2: If the second mark is located within the first mark and the second mark is used to delineate content that does not need to be identified, the electronic device can use the other areas in the first region delineated by the first mark, excluding the second region delineated by the second mark, as the third region, and then identify the target content from the screen content of the electronic device based on the third region.
[0317] The third area is used to indicate the content that needs to be identified. The operation of the electronic device identifying the target content from the screen content of the electronic device based on the third area is similar to the operation of the electronic device identifying the target content from the screen content of the electronic device based on the first area in the first method described above, and will not be described again in this embodiment.
[0318] For example, if the content located in the third area of the screen content of the electronic device is text content, the electronic device determines that the text paragraph containing the text content in the third area is the third paragraph. The electronic device determines the overlap between each sentence in the third paragraph and the third area based on the coordinates of each sentence in the third paragraph and the coordinate range of the third area. Based on the overlap between each sentence in the third paragraph and the third area, the electronic device identifies the target sentence from the third paragraph, which is the target content. Alternatively, the electronic device determines the overlap between each character in the third paragraph and the third area based on the coordinates of each character in the third paragraph and the coordinate range of the third area. Based on the overlap between each character in the third paragraph and the third area, the electronic device identifies the target character from the third paragraph, which is the target content.
[0319] If the content on the screen of the electronic device located in the third area is table content, the electronic device determines that the table content in the third area belongs to the third table. The electronic device determines the overlap between each cell in the third table and the third area based on the coordinates of each cell in the third table and the coordinate range of the third area; based on the overlap between each cell in the third table and the third area, it identifies the text in the target cell from the third table, and the text in the target cell is the target content.
[0320] If the content located in the third area of the screen content of the electronic device is image content, the electronic device determines that the image content in the third area belongs to the third image, identifies the third image based on the coordinates of the third image, and the third image is the target content.
[0321] If the content in the third area of the screen of the electronic device is formula content, then the electronic device identifies the formula (which can be called the third formula) in the third area, and the third formula is the target content.
[0322] In some embodiments, if the second marker is located within the first marker and is used to delineate content that does not need to be identified, the electronic device can identify the target content from the screen content of the electronic device using method 2 described above. Alternatively, the electronic device can retain the first marker and discard the second marker within the specified markers, and then identify the target content from the screen content of the electronic device using the first method described above. Alternatively, the electronic device can discard the first marker and retain the second marker within the specified markers, and then identify the target content from the screen content of the electronic device using the second method described above.
[0323] For example, such as Figure 25 As shown, the electronic device acquires a specified tag, assuming the specified tag is as follows: Figure 26 Figure (a) shows a first mark and a second mark. The electronic device then determines the overlap between the first mark and the second mark, because... Figure 26 In Figure (a), the second mark is inside the first mark, so the electronic device can determine the overlap between the first and second marks as follows: the second mark is inside the first mark. In this case, the electronic device can retain both the first and second marks, thus obtaining... Figure 26 The designated marker shown in Figure (b) allows for the identification of target content from the screen content of the electronic device using method 2 described above. Alternatively, the electronic device can retain the first marker and discard the second marker, resulting in... Figure 26 The specified marker shown in Figure (c) includes the first marker but excludes the second marker, thus allowing the target content to be identified from the screen content of the electronic device using the first method described above. Alternatively, the electronic device can discard the first marker and retain the second marker, obtaining... Figure 26 The specified marker shown in Figure (d) includes the second marker but does not include the first marker, so the target content can be identified from the screen content of the electronic device in the second way described above.
[0324] Method 3: If the second mark is located within the first mark and the second mark is used to cover content that does not need to be recognized, the electronic device recognizes the target content from the screen content of the electronic device according to the first area defined by the first mark, and then deletes the content covered by the second mark from the recognized target content.
[0325] The operation of the electronic device identifying target content from the screen content of the electronic device based on the first area defined by the first mark is similar to the operation of the electronic device identifying target content from the screen content of the electronic device based on the first area defined by the first mark in the first method described above, and will not be described again in this embodiment.
[0326] The operation of the electronic device to delete the content occluded by the second marker from the identified target content can be as follows: If the identified target content is text, the electronic device determines the occlusion degree of each sentence in the target content based on the coordinates of each sentence in the target content and the coordinates of the second marker. If the occlusion degree of a sentence in the target content is greater than or equal to the first occlusion degree, then the sentence is deleted from the target content; or, the electronic device determines the occlusion degree of each character in the target content based on the coordinates of each character in the target content and the coordinates of the second marker. If the occlusion degree of a character in the target content is greater than or equal to the first occlusion degree, then the character is deleted from the target content.
[0327] If the identified target content is a table, the electronic device determines the occlusion degree of each cell in the target content based on the coordinates of each cell in the target content and the coordinates of the second marker. If the occlusion degree of a cell in the target content is greater than or equal to the second occlusion degree, the text in that cell is deleted from the target content.
[0328] Specifically, if the identified target content is an image or formula, the electronic device discards the second marker, meaning that the electronic device does not need to delete the content obscured by the second marker from the identified target content.
[0329] In some embodiments, if the second marker is located within the first marker and is used to obscure content that does not need to be identified, the electronic device can identify the target content from the screen content of the electronic device using method 3 described above. Alternatively, the electronic device can retain the first marker and discard the second marker within the specified marker, and then identify the target content from the screen content of the electronic device using the first method described above. Alternatively, the electronic device can discard the first marker and retain the second marker within the specified marker, and then identify the target content from the screen content of the electronic device using the second method described above.
[0330] Method 4: If the first mark is located within the second mark, and the second mark is used to delineate content that does not need to be identified, then the electronic device identifies the target content from the screen content of the electronic device based on the first area delineated by the first mark, and identifies the target content from the screen content of the electronic device based on the second area delineated by the second mark.
[0331] The operation of the electronic device identifying target content from the screen content of the electronic device based on the first area defined by the first mark is similar to the operation of the electronic device identifying target content from the screen content of the electronic device based on the first area defined by the first mark in the first method described above, and will not be described again in this embodiment.
[0332] The operation of the electronic device identifying target content from the screen content of the electronic device based on the second area defined by the second mark is similar to the operation of the electronic device identifying target content from the screen content of the electronic device based on the second area defined by the second mark in the second method described above, and will not be described again in this embodiment.
[0333] In some embodiments, if the first marker is located within the second marker and the second marker is used to delineate content that does not need to be identified, the electronic device can identify the target content from the screen content of the electronic device using method 4 described above. Alternatively, the electronic device can retain the first marker and discard the second marker within the specified markers, and then identify the target content from the screen content of the electronic device using the first method described above. Alternatively, the electronic device can discard the first marker and retain the second marker within the specified markers, and then identify the target content from the screen content of the electronic device using the second method described above.
[0334] For example, such as Figure 27 As shown, the electronic device acquires a specified tag, assuming the specified tag is as follows: Figure 28 Figure (a) shows a first mark and a second mark. The electronic device then determines the overlap between the first mark and the second mark, because... Figure 28 In Figure (a), the first mark is inside the second mark, so the electronic device can determine that the overlap between the first and second marks is such that the first mark is inside the second mark. In this case, the electronic device can retain both the first and second marks, thus obtaining... Figure 28 The designated marker shown in Figure (b) allows for the identification of target content from the screen content of the electronic device using method 4 described above. Alternatively, the electronic device can retain the first marker and discard the second marker, resulting in... Figure 28 The specified marker shown in Figure (c) includes the first marker but excludes the second marker, thus allowing the target content to be identified from the screen content of the electronic device using the first method described above. Alternatively, the electronic device can discard the first marker and retain the second marker, obtaining... Figure 28 The specified marker shown in Figure (d) includes the second marker but does not include the first marker, so the target content can be identified from the screen content of the electronic device in the second way described above.
[0335] Method 5: If the first and second markers intersect and the second marker is used to delineate content that does not need to be identified, the electronic device retains the first marker and uses the intersection of the first and second markers as the new second marker; or, it retains the second marker and uses the union of the second and first markers as the new first marker. Then, the target content is identified from the screen content of the electronic device using Method 2 described above.
[0336] For example, such as Figure 29 As shown, the electronic device acquires a specified tag, assuming the specified tag is as follows: Figure 30 Figure (a) shows a first mark and a second mark. The electronic device then determines the overlap between the first mark and the second mark, because... Figure 30 In Figure (a), the first mark intersects with the second mark, so the electronic device can determine that the overlap between the first and second marks is: the first mark intersects with the second mark. In this case, the electronic device can retain the first mark and take the intersection of the first and second marks as the new second mark, thus obtaining... Figure 30 The designated marker shown in Figure (b) is located within the first marker, thus allowing the target content to be identified from the screen content of the electronic device using method 2 described above. Alternatively, the electronic device can retain the second marker and use the union of the second marker and the first marker as the new first marker, obtaining... Figure 30The designated mark shown in Figure (c) is located within the first mark, thus the target content can be identified from the screen content of the electronic device through the above method 2.
[0337] Method 6: If the first mark and the second mark intersect and the second mark is used to cover content that does not need to be identified, the electronic device retains the first mark and uses the part of the second mark that is inside the first mark as the new second mark, and then identifies the target content from the screen content of the electronic device through Method 3 above.
[0338] For example, such as Figure 31 As shown, the electronic device acquires a specified tag, assuming the specified tag is as follows: Figure 32 Figure (a) shows a first mark and a second mark. The electronic device then determines the overlap between the first mark and the second mark, because... Figure 32 As shown in Figure (a), the first mark intersects with the second mark, so the electronic device can determine that the overlap between the first and second marks is: the first mark intersects with the second mark. In this case, the electronic device can retain the first mark and take the portion of the second mark that lies within the first mark as the new second mark, thus obtaining... Figure 32 The designated mark shown in Figure (b) is located within the first mark, thus the target content can be identified from the screen content of the electronic device through the above method 3.
[0339] In some embodiments, if the screen content of an electronic device is a video playback interface, the screen content recognition method provided in this application can be used to recognize the video image of a video in a paused state within the video playback interface. In this case, after the user marks a specified marker on the video image in the video playback interface, they can also choose to apply the marked specified marker in real time during video playback.
[0340] For example, an electronic device may provide a designated button. By operating this designated button, the user can choose to continuously apply the designated marker during video playback. In this case, the electronic device can identify the target content from each frame of the video based on the drawn designated marker. Alternatively, the user can choose to continuously apply the designated marker while the video playback time falls within certain time periods. In this case, when the video playback reaches a point within these time periods, the electronic device can identify the target content from each frame of the video based on the drawn designated marker.
[0341] Optionally, after the electronic device identifies target content during video playback, it can store the identified target content to the system clipboard when the video playback ends. Afterward, the electronic device can automatically paste the content from the system clipboard into the edit box, or the user can manually paste the content from the system clipboard into the edit box.
[0342] Optionally, after the electronic device identifies the target content during video playback, it can also store the identified target content to a specified file when the video playback ends, so that users can easily view the identified target content in the specified file after the video playback ends.
[0343] For example, such as Figure 33 As shown, the electronic device receives a video stream and then renders it onto a visual interface for video playback. Let's assume this visual interface is... Figure 34 The video playback interface 3401 is shown below. In this case, if the user wants to obtain the subtitle information of the video played in the video playback interface 3401, the user can pause a certain frame of the video image and then draw a specified mark on the screen, for example, as shown. Figure 34 As shown in Figure (a), the user can draw the first mark in the subtitle area of the video playback interface 3401, or, as ... Figure 34 As shown in Figure (b), the user can draw a second mark in the non-subtitle area of the video playback interface 3401. In this way, the electronic device can identify target content from each frame of the video image during playback based on the specified mark, for example, in... Figure 34 This refers to real-time subtitle recognition. In this case, the electronic device can output the recognized subtitle information after the video playback ends, allowing users to easily and conveniently obtain the subtitle information for the entire video, thus improving the user experience.
[0344] To facilitate understanding, the following will be combined with... Figure 35 The above screen content recognition method will be illustrated by example.
[0345] Figure 35 This is a schematic diagram of a screen content recognition method provided in an embodiment of this application. See also... Figure 35 The method includes the following steps (1) to (8):
[0346] (1) The user draws a trajectory line on the screen of the electronic device.
[0347] (2) The electronic device obtains the specified marker (i.e., the first marker and / or the second marker) based on the drawn trajectory line.
[0348] (3) If the specified mark includes only the first mark or the second mark, the electronic device will not process the specified mark. If the specified mark includes both the first mark and the second mark, and the first mark and the second mark intersect, then the above applies. Figure 12 In step 1203 of the embodiment, method 5 or method 6 in the third method performs corresponding processing on the specified mark. If the specified mark includes a first mark and a second mark, and the second mark is located within the first mark, then it is processed according to the above. Figure 12 In step 1203 of the embodiment, method 2 or method 3 in the third method performs corresponding processing on the specified mark. If the specified mark includes a first mark and a second mark, and the first mark is located within the second mark, the electronic device processes the specified mark according to the above description. Figure 12 In step 1203 of the embodiment, method 4 of the third method performs corresponding processing on the specified mark. If the specified mark includes a first mark and a second mark, and the first mark and the second mark do not overlap, then it is processed according to the above. Figure 12 In step 1203 of the embodiment, method 1 of the third method processes the specified mark accordingly. Then, the electronic device performs subsequent steps based on the specified mark to identify target content from the screen content of the electronic device.
[0349] (4) Electronic devices perform text line detection, component area detection, and component type detection on the screen content to determine the text area, table area, image area, and formula area in the screen content.
[0350] The components can include tables, images, and formulas. For determining the text region, after detecting text lines, text paragraph analysis can be performed to determine the coordinates of each text paragraph and the coordinates of each text line within each paragraph.
[0351] It is worth noting that in this embodiment, lightweight detection can be achieved through text line detection, component area detection, and component type detection. Since this detection does not identify specific content, it can avoid the situation where the entire document is slow to be recognized.
[0352] (5) If the content indicated by the specified mark is in a text area, the electronic device shall follow the above. Figure 12 The first case in step 1203 of the embodiment is used to identify target content. If the content indicated by the specified marker is in the table area, the electronic device follows the above... Figure 12 The second scenario in step 1203 of the embodiment is used to identify the target content. If the content indicated by the specified marker is within the image area, the electronic device proceeds according to the above... Figure 12 The third scenario in step 1203 of the embodiment is used to identify the target content. If the content indicated by the specified marker is in the formula area, the electronic device follows the above... Figure 12The fourth case in step 1203 of the embodiment is used to identify the target content.
[0353] (6) If the identified target content has multiple types, such as at least two of text, tables, pictures and formulas, the electronic device can sort the multiple target contents according to the corresponding types.
[0354] (7) The electronic device stores the target content to the system clipboard.
[0355] (8) The electronic device pastes the target content from the system clipboard into the editing box.
[0356] The following is combined Figure 36 To match the above text Figure 35 The embodiments are illustrated in detail.
[0357] Figure 36 This is a flowchart of a screen content recognition method provided in an embodiment of this application. See also... Figure 36 The method includes the following steps 3601 to 3610:
[0358] Step 3601: The user draws a trajectory line on the screen and specifies the marker acquisition module to obtain the drawn trajectory line.
[0359] Step 3602: The specified marker acquisition module acquires the specified marker based on the drawn trajectory line.
[0360] Step 3603: The designated marker processing module processes the designated marker based on the overlap between the first and second markers. See above for details. Figure 35 Step (3) in the embodiment.
[0361] Step 3604: If the content indicated by the specified marker is text content, the content recognition module performs text line detection, then performs paragraph analysis to obtain the coordinates of each text paragraph and the coordinates of each text line within each text paragraph.
[0362] Step 3605: The content recognition module determines candidate paragraphs based on the coordinate range of the text content indicated by the specified markers and the coordinates of each text paragraph. For details, please refer to the above. Figure 12 The first case in step 1203 of the embodiment.
[0363] The coordinate range of the text content indicated by the specified marker is the coordinate range of the first region enclosed by the first marker, and / or the coordinate range of the second region enclosed by the second marker or the coordinates of the second marker.
[0364] Candidate paragraphs are text paragraphs containing text content within the first region defined by the first marker, and / or, candidate paragraphs are text paragraphs containing text content other than text content within the second region defined by the second marker or text content obscured by the second marker.
[0365] As an example, the target sentence can be identified through steps 3606 and 3607 below. As another example, the target character can be identified through step 3608 below.
[0366] Step 3606: The content recognition module obtains the coordinates of each sentence in the candidate paragraph through the NLP module.
[0367] Step 3607: The content recognition module identifies the target sentence based on the coordinate range of the text content indicated by the specified marker and the coordinates of each sentence in the candidate paragraph.
[0368] Step 3608: The content recognition module identifies the target character based on the coordinate range of the text content indicated by the specified marker and the coordinates of each character in the candidate paragraph.
[0369] Step 3609: The content recognition module stores the recognized content to the system clipboard.
[0370] Step 3610: Paste the contents of the system clipboard into the edit box when the edit box gains focus.
[0371] The following is combined Figure 37 This section explains the content recognition process when the content indicated by the specified marker is text content.
[0372] See Figure 37 After performing text line detection and text paragraph analysis, the coordinates of the text paragraphs and the coordinates of the text lines within the paragraphs are obtained. Then, candidate paragraphs are determined based on specified markers, the coordinates of the text paragraphs, and the coordinates of the text lines within the paragraphs. For specific instructions, please refer to the above text. Figure 36 Step 3605 in the embodiment. Afterwards, it can be determined whether the specified marker is the first marker or the second marker.
[0373] If the specified marker is the first marker, text recognition is performed on the candidate paragraphs. Then, sentence coordinate detection is performed on the candidate paragraphs to obtain the coordinates of each sentence. The overlap of each sentence is determined based on its coordinates; if a sentence has a high overlap, it is extracted; otherwise, it is not extracted. Alternatively, character coordinate detection is performed on the candidate paragraphs to obtain the coordinates of each character. The overlap of each character is determined based on its coordinates; if a character has a high overlap, it is extracted; otherwise, it is not extracted.
[0374] If the specified marker is the second marker, text recognition is performed on the candidate paragraphs. Then, sentence coordinate detection is performed on the candidate paragraphs to obtain the coordinates of each sentence. The overlap of each sentence is determined based on its coordinates; if a sentence has a high overlap, it is not extracted; if a sentence has a low overlap, it is extracted. Alternatively, character coordinate detection is performed on the candidate paragraphs to obtain the coordinates of each character. The overlap of each character is determined based on its coordinates; if a character has a high overlap, it is not extracted; if a character has a low overlap, it is extracted.
[0375] The following is combined Figure 38 This section explains the content recognition process when the content indicated by the specified marker is table content.
[0376] See Figure 38 After obtaining the specified tag, if the content indicated by the specified tag is within a table area, then table structure detection is performed. Afterwards, it can be determined whether the specified tag is the first tag or the second tag.
[0377] If the specified marker is the first marker, the target cell is determined based on the overlap of cells in the table. The target cell is aligned vertically and horizontally to obtain a sub-table of row a and column b. Text recognition is performed on the sub-table, and then the sub-table is output.
[0378] If the specified marker is the second marker, text recognition is performed on the entire table. The target cell is determined based on the overlap of cells in the table. The contents of all cells in the entire table except the target cell are cleared, and then the entire table is output.
[0379] The following is combined Figure 39 This section explains the content recognition process when the content indicated by the specified marker is image content.
[0380] See Figure 39 After obtaining the specified marker, if the content indicated by the marker is within the image area, then the marker is determined to be either the first marker or the second marker. If the marker is the second marker, the image is not recognized. If the marker is the first marker, the image is recognized, corrected, and the recognized image is output.
[0381] The following is combined Figure 40 This section explains the content recognition process when the content indicated by the specified marker is image content.
[0382] See Figure 40After obtaining the specified marker, if the content indicated by the specified marker is within the formula area, then determine whether the specified marker is the first marker or the second marker. If the specified marker is the second marker, the formula is not recognized. If the specified marker is the first marker, the formula is recognized and output.
[0383] Figure 41 This is a schematic diagram of the structure of a screen content recognition device provided in an embodiment of this application. The device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figures 1 to 2 The electronic device 100 described in the embodiment. See also... Figure 41 The device includes: a first acquisition module 4101, a second acquisition module 4102, and an identification module 4103.
[0384] The first acquisition module 4101 is used to execute the above. Figure 12 Step 1201 in the embodiment;
[0385] The second acquisition module 4102 is used to execute the above. Figure 12 Step 1202 in the embodiment;
[0386] The recognition module 4103 is used to perform the above. Figure 12 Step 1203 in the embodiment.
[0387] Optionally, the screen content is an image or interface displayed by the electronic device, such as an application interface, a video playback interface, or a camera preview interface.
[0388] Optionally, the first mark is a closed shape and the second mark is an open shape; or, the first mark is a first preset shape and the second mark is a second preset shape, and the shapes of the first preset shape and the second preset shape are different; or, the first mark is a closed shape and the second mark is a closed shape, and the line thicknesses of the first mark and the second mark are different; or, the first mark is a closed shape and the second mark is a closed shape, and the line colors of the first mark and the second mark are different.
[0389] Optionally, the first marker is a closed shape, the second marker is a non-closed shape, and the second acquisition module 4102 is used for:
[0390] Perform polygon fitting on n trajectory lines;
[0391] If a polygon is fitted by at least one of the n trajectory lines, then the fitted polygon is determined as the first label, and the trajectory lines among the n trajectory lines that do not fit a polygon are determined as the second label.
[0392] If multiple polygons are fitted through at least one of the n trajectory lines, then the first mark is determined based on the fitted multiple polygons and the overlap of the multiple polygons, and the trajectory lines among the n trajectory lines that do not fit polygons are determined as the second mark.
[0393] If no polygon can be fitted using any of the n trajectory lines, then the n trajectory lines are determined as the second marker.
[0394] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to delineate the content that does not need to be identified or to obscure the content that does not need to be identified. The identification module 4103 is used for:
[0395] If the specified marker includes the first marker but does not include the second marker, then the target content is identified from the screen content of the electronic device based on the first area defined by the first marker;
[0396] If the specified marker includes the second marker but does not include the first marker, the target content is identified from the screen content of the electronic device based on the second area defined by the second marker or based on the content obscured by the second marker.
[0397] Optionally, the identification module 4103 is used for:
[0398] If the content located in the first area of the screen content of the electronic device is text content, then the text paragraph containing the text content in the first area is determined to be the first paragraph;
[0399] Based on the coordinates of each sentence in the first paragraph and the coordinate range of the first region, determine the degree of overlap between each sentence in the first paragraph and the first region. Based on the degree of overlap between each sentence in the first paragraph and the first region, identify the target sentence from the first paragraph. The target sentence is the target content. Alternatively, based on the coordinates of each character in the first paragraph and the coordinate range of the first region, determine the degree of overlap between each character in the first paragraph and the first region. Based on the degree of overlap between each character in the first paragraph and the first region, identify the target character from the first paragraph. The target character is the target content.
[0400] Optionally, the identification module 4103 is used for:
[0401] If the content in the first area of the screen content of the electronic device is table content, then the table to which the table content in the first area belongs is determined to be the first table;
[0402] Based on the coordinates of each cell in the first table and the coordinate range of the first region, determine the degree of overlap between each cell in the first table and the first region;
[0403] The text in the target cell is identified from the first table based on the degree of overlap between each cell in the first table and the first area. The text in the target cell is the target content.
[0404] Optionally, the identification module 4103 is used for:
[0405] If the content located in the first area of the screen content of the electronic device is image content, then the image to which the image content in the first area belongs is determined to be the first image;
[0406] The first image is identified based on its coordinates, and the first image is the target content.
[0407] Optionally, the identification module 4103 is used for:
[0408] If the content in the first area of the electronic device's screen is formula content, then the formula containing the formula content in the first area is identified, and the formula containing the formula content in the first area is the target content.
[0409] Optionally, the first marker is used to delineate the content to be identified, and the second marker is used to delineate the content that does not need to be identified or to obscure the content that does not need to be identified. The identification module 4103 is used for:
[0410] If the specified markers include a first marker and a second marker, and the first marker and the second marker do not overlap, then the first marker is retained and the second marker is discarded in the specified markers, and the target content is identified from the screen content of the electronic device based on the first area defined by the first marker; or,
[0411] If the specified markers include a first marker and a second marker, and the first marker and the second marker do not overlap, then the first marker is discarded and the second marker is retained in the specified markers. The target content is identified from the screen content of the electronic device based on the second area enclosed by the second marker or based on the content occluded by the second marker.
[0412] Optionally, the first marker is used to delineate the content to be recognized, and the second marker is used to delineate the content that does not need to be recognized. The recognition module 4103 is used for:
[0413] If the specified markers include a first marker and a second marker, and the second marker is located within the first marker, then the other areas in the first region enclosed by the first marker, excluding the second region enclosed by the second marker, are determined as the third region, and the target content is identified from the screen content of the electronic device based on the third region.
[0414] Optionally, the first marker is used to delineate the content to be recognized, and the second marker is used to obscure the content that does not need to be recognized. The recognition module 4103 is used for:
[0415] If the specified marker includes a first marker and a second marker, and the second marker is located within the first marker, then the target content is identified from the screen content of the electronic device based on the first area defined by the first marker, and the content obscured by the second marker is deleted from the identified target content.
[0416] Optionally, the first marker is used to delineate the content to be recognized, and the second marker is used to delineate the content that does not need to be recognized. The recognition module 4103 is used for:
[0417] If the specified marker includes a first marker and a second marker, and the first marker and the second marker intersect, then the first marker is retained and the intersection of the first marker and the second marker is taken as the new second marker; or, the second marker is retained and the union of the second marker and the first marker is taken as the new first marker.
[0418] The area outside the second area defined by the second mark within the first region is defined as the third region, and the target content is identified from the screen content of the electronic device based on the third region.
[0419] Optionally, the first marker is used to delineate the content to be recognized, and the second marker is used to obscure the content that does not need to be recognized. The recognition module 4103 is used for:
[0420] If the specified mark includes a first mark and a second mark, and the first mark and the second mark intersect, then the first mark is retained, and the portion of the second mark that is inside the first mark is taken as the new second mark;
[0421] Based on the first area defined by the first mark, target content is identified from the screen content of the electronic device, and content obscured by the second mark is deleted from the identified target content.
[0422] Optionally, the device further includes:
[0423] The storage module is used to store the identified target content to the system clipboard;
[0424] The paste module is used to paste the target content stored in the system clipboard into the edit box if any edit box is detected to have focus.
[0425] In this embodiment, it is not necessary to recognize the entire screen content of the electronic device. Instead, target content that meets the user's needs can be identified from the screen content based on the user's specified markers. The recognition time is short, content recognition can be completed quickly, and power consumption can be saved. Furthermore, since the user can select the content to be recognized or not recognized according to their own needs—that is, the user can select a local area of content on the screen for recognition or not, or select multiple discontinuous areas of content on the screen for recognition or not—this application can achieve cross-line recognition of screen content, thereby making screen content recognition more flexible.
[0426] It should be noted that the screen content recognition device provided in the above embodiments is only illustrated by the division of the above functional modules when recognizing screen content. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0427] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.
[0428] The screen content recognition device and screen content recognition method provided in the above embodiments belong to the same concept. The specific working process and technical effects of the units and modules in the above embodiments can be found in the method embodiments section, and will not be repeated here.
[0429] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).
[0430] The above-described embodiments are optional embodiments provided by this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the technical scope disclosed in this application should be included within the protection scope of this application.
Claims
1. A method of screen content recognition, the method comprising: The method is applied to an electronic device, and the method comprises: displaying a first interface; receiving a first operation on the first interface, and displaying a first figure on the first interface.
2. The method of claim 1, wherein, The method further comprises: displaying a second interface before displaying the first interface; receiving a second operation while displaying the second interface, and displaying the first interface; receiving a third operation while displaying the first interface, and displaying a third interface.
3. The method of claim 2, wherein, The first interface comprises first content, and the first figure indicates the first content.
4. The method of any one of claims 1-3, wherein, The method further comprises: displaying a second figure on the first interface after displaying the first figure on the first interface.
5. The method of claim 4, wherein, A corresponding display region of the first figure on the first interface is a first display region, and a corresponding display region of the second figure on the first interface is a second display region, and the first display region is different from the second display region.
6. The method of claim 5, wherein, The first interface further comprises second content, and the second figure indicates the second content.
7. The method of any one of claims 1-3, wherein, The method further comprises: receiving a fourth operation on the first figure while displaying the first figure on the first interface, and displaying a third figure on the first interface, wherein a corresponding display region of the first figure on the first interface is a first display region, and a corresponding display region of the third figure on the first interface is a third display region, and the third display region is different from the first display region.
8. The method of claim 7, wherein, The third figure is a figure obtained by adjusting the first figure.
9. The method of any one of claims 1-3, wherein, The method further comprises: displaying a fourth interface after displaying the first figure on the first interface, and the fourth interface comprises a first editing box; receiving a fifth operation, and displaying first content in the first editing box.
10. The method of any one of claims 1-9, wherein, The first interface comprises third content, and the third content comprises a first character and a second character, and the method further comprises: receiving a sixth operation on the first interface, and displaying a fourth figure on the first interface, wherein the fourth figure is a closed figure, a corresponding display region of the fourth figure on the first interface is a fourth display region, a corresponding display region of the first character does not overlap with the fourth display region, and a corresponding display region of the second character at least partially overlaps with the fourth display region.
11. The method of claim 10, wherein, The method further comprises: displaying a fifth interface after displaying the fourth figure on the first interface, and the fifth interface comprises a second editing box; receiving a seventh operation, and displaying fourth content in the second editing box, wherein the fourth content comprises the second character, and the fourth content does not comprise the first character.
12. The method of any one of claims 1-11, wherein, The first interface comprises fifth content, and the fifth content comprises a first text and a second text, a line where the first text is located is a first line, and a line where the second text is located is a second line, and the method further comprises: The eighth operation is received on the first interface, a fifth graphic is displayed on the first interface, the fifth graphic is a closed graphic, a corresponding display area of the fifth graphic on the first interface is a fifth display area, the first row corresponding display area does not overlap with the fifth display area, and the second row corresponding display area at least partially overlaps with the fifth display area.
13. The method of claim 12, wherein, The method further includes: After the fifth graphic is displayed on the first interface, a sixth interface is displayed, and the sixth interface includes a third edit box. The ninth operation is received, and sixth content is displayed in the third edit box, the sixth content includes the second text, and the sixth content does not include the first text.
14. The method of claims 1-13, wherein, The first interface includes a first picture, and the method further includes: The tenth operation is received on the first interface, a sixth graphic is displayed on the first interface, a corresponding display area of the sixth graphic on the first interface is a sixth display area, and a first part of the first picture is displayed in the sixth display area.
15. The method of claim 14, wherein, The method further includes: After the sixth graphic is displayed on the first interface, a seventh interface is displayed, and the seventh interface includes a fourth edit box. The eleventh operation is received, and seventh content is displayed in the fourth edit box, the seventh content includes a second part of the first picture, and the second part includes the first part.
16. An electronic device, comprising: An electronic device includes a memory and one or more processors; the memory is coupled to the one or more processors, the memory is configured to store computer program code, the computer program code includes computer instructions, and the one or more processors invoke the computer instructions to enable the electronic device to perform the method of any one of claims 1-15.
17. A computer program product, characterised in that, A computer program is included, and when the computer program is executed by a processor, the method of any one of claims 1-15 is implemented.
18. A computer-readable storage medium comprising instructions, wherein, When the instructions are executed on an electronic device, the electronic device performs the method of any one of claims 1-15.