Screen identification method and electronic equipment
By implementing the screen recognition method on an electronic device, identifying and marking content in the user interface, the problem of low information acquisition efficiency in the prior art is solved, and the human-computer interaction efficiency is improved.
Patent Information
- Application Number
- CN202311821050.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2023-12-26
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively identify and mark content in the user interface, making it difficult for users to obtain information efficiently.
By implementing the screen recognition method on an electronic device, the content in the user interface is recognized, and the corresponding text and graphic code are marked based on the function of recommending matching the interface according to the recognition result.
It improves human-computer interaction efficiency, allows users to quickly obtain information in the user interface, and enhances the intelligent recognition capabilities of electronic devices.
Smart Images

Figure CN119987601A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 3, 2023, with application number 202311461131.5 and invention name “A method and electronic device for intelligent identification”, the entire contents of which are incorporated by reference in this application. Technical Field
[0002] The embodiments of the present application relate to the field of terminal technology, and in particular to a screen recognition method and an electronic device. Background Art
[0003] In the daily use of electronic devices such as mobile phones and tablets, the mobile phone may need to recognize the content in the user interface, such as identifying products and texts in the user interface.
[0004] In the prior art, although there are some solutions for recognizing the content in the user interface, such as recognizing the text in the user interface through the optical character recognition (OCR) technology, the recognition results cannot be presented well, and thus cannot help users to obtain information efficiently. Summary of the invention
[0005] The present application provides a screen recognition method and an electronic device, which can identify the content in the user interface and recommend functions matching the interface based on the recognition results, so that the user can perform corresponding processing on the content in the user interface, thereby improving the efficiency of human-computer interaction.
[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, the present application provides a screen recognition method, which is applied to an electronic device. A first interface is displayed, and the first interface includes a first object entity, a first text entity, and a second text entity. In response to a first trigger operation on the first interface, a first text mark is added to the first text entity, and no mark is added to the second text entity. A second interface is displayed, and the second interface includes a second object entity, a third text entity, and a fourth text entity. In response to a second trigger operation on the second interface, a second text mark is added to the third text entity, and no mark is added to the fourth text entity.
[0008] Among them, the entity categories of the first item entity and the second item entity are different, such as the first item entity is a building, and the second item entity is an animal. That is to say, the first interface and the second interface include item entities of different entity categories. The entity categories of the first text entity and the fourth text entity are the same, such as both are telephone numbers, and the entity categories of the second text entity and the third text entity are the same, such as both are addresses. The entity categories of the first text entity and the second text entity are different. That is, not all text entities of entity categories will be marked. Moreover, text entities of the same entity category are not marked in both the first interface and the second interface.
[0009] In summary, by using the present application, an electronic device can mark text entities of the first category (entity category of the first text entity and the fourth text entity) in an interface with a certain category (entity category of the first item entity), but not mark text entities of the second category (entity category of the second text entity and the third text entity); and an electronic device can mark text entities of the second category in an interface with another category (entity category of the second item entity), but not mark text entities of the first category. It can be seen that the electronic device can mark matching text entities based on the item entity included in the user interface, thereby providing users with marks for quickly obtaining information from the user interface, which is conducive to improving the efficiency of human-computer interaction.
[0010] In a possible design of the first aspect, the entity categories of the text entity include at least two of the following: address, telephone number, flight information, express delivery number, email address, website link, certificate number for identity identification, and graphic code. The entity categories of the object entity include at least two of the following: animals, plants, buildings, and food.
[0011] That is, the electronic device can mark the above-mentioned text entities that match the animals, plants, buildings and foods in the user interface.
[0012] In a possible design of the first aspect, the first text mark indicates the entity category of the first text entity, such as the first text mark is a category icon of the entity category. For example, if the first text entity is a phone number, the first mark can be a phone icon.
[0013] Alternatively, the first text entity is associated with multiple services. Taking the first text entity as a phone number as an example, the phone number can be associated with multiple services such as making calls and adding to contacts. The first text tag can indicate the first service that the user is most interested in among multiple services, such as the first text tag being a service icon for the first service. Among them, the electronic device can regard the service that the user has selected the most times under the text entity of the first category as the first service of greatest interest, and the first category is the entity category of the first text entity. In this way, the electronic device can mark the text entity with a tag corresponding to the service that the user is interested in (such as a service icon).
[0014] It should be understood that the specific content of the second text mark can also refer to the first text entity, which will not be repeated here.
[0015] In a possible design of the first aspect, after adding a first text tag to a first text entity in response to a first trigger operation on a first interface, the method further includes: displaying a third interface in response to a third trigger operation on the first text entity or the first text tag, wherein the third interface includes a plurality of service options, and the plurality of service options correspond one to one to a plurality of services. Among the plurality of service options, the service option of the first service is displayed first.
[0016] That is, the electronic device ranks the option of the first service that the user is most interested in first among the multiple service options so that the user can use the first service through the option of the first service. Of course, for other unmarked text entities of the first category, the electronic device can also respond in the same way, that is, display the service option of the first service first.
[0017] It should be understood that in response to the third trigger operation on the third text entity or the second text mark, the response of the electronic device is also the same, which will not be elaborated here.
[0018] In a possible design of the first aspect, the first interface also includes a fifth text entity. The above method also includes: in response to a first trigger operation on the first interface, and the entity category of the fifth text entity is different from that of the first text entity, adding a third text mark to the fifth text entity. If the entity category of the fifth text entity is the same as that of the first text entity, no mark is added to the fifth text.
[0019] That is to say, for text entities of the same entity category, the electronic device only displays a mark for one of the text entities to avoid repeated marks for text entities of the same entity category.
[0020] In a possible design of the first aspect, the first interface also includes a sixth text entity. The above method also includes: in response to a first trigger operation on the first interface, and the fourth entity mark of the sixth text entity and the first text mark are not blocked, adding a fourth text mark to the sixth text entity. If the fourth text mark and the first text mark are blocked, no mark is added to the sixth text.
[0021] That is to say, the electronic device will display all the markers only when the markers do not block each other.
[0022] In a possible design of the first aspect, the first interface further includes a third item entity. The method further includes: in response to a first trigger operation on the first interface, highlighting the first item entity, and displaying a first quick entry around the first item entity. In response to a fourth trigger operation on the third item entity, highlighting the third item entity, and displaying a second quick entry around the first item entity. The highlighted item entity can be understood as a focus item entity.
[0023] That is, in response to a trigger operation by the user, the electronic device may switch the focus item entity and display a shortcut entry corresponding to the focus item entity so that the user can obtain information about the focus item entity.
[0024] In a possible design of the first aspect, the above-mentioned response to the first trigger operation on the first interface, highlighting the first item entity, and displaying the first quick entry around the first item entity, includes: responding to the first trigger operation on the first interface, and the first item entity meets the first condition, highlighting the first item entity, and displaying the first quick entry around the first item entity.
[0025] The first condition includes at least one of the following:
[0026] Condition 1: The area of the first item entity is greater than the area of the third item entity. For example, the area of the first item entity is the largest item entity in the first interface. That is, the electronic device may preferentially use the item entity with the larger area as the focus item entity.
[0027] Condition 2: The area blocked by the first item entity is smaller than the area blocked by the third item entity. It can be understood that the smaller the area blocked by the first item entity is, the higher the integrity of the first item entity is and the more comprehensive the display is. In this way, the electronic device can prioritize the item entity that can be fully displayed as the focus item entity.
[0028] Condition 3: The clarity of the edge line of the first object entity is higher than the clarity of the edge line of the third object entity. The higher the clarity of the edge of the first object entity, the more accurately the electronic device can cut out the first object entity. In other words, the electronic device can give priority to the object entity with more accurate cutting out as the focus object entity.
[0029] At this point, it should be noted that the combination of the above conditions 1, 2 and 3, that is, the area of the first item entity is larger than the area of the third item entity, the area blocked by the first item entity is smaller than the area blocked by the third item entity, and the clarity of the edge line of the first item entity is higher than the clarity of the edge line of the third item entity, indicates that the first item entity has a larger area, higher integrity and clearer boundaries. In other words, the electronic device can give priority to the item entity with a larger area and more accurate and complete cutout as the focus item entity.
[0030] In a possible design of the first aspect, the method further includes: in response to a first trigger operation on the first interface, displaying a first item mark on the third item entity, such as the first item mark being a circle pattern. The fourth trigger operation includes a trigger operation on the first item mark. That is, the electronic device can use the first item mark as a clear trigger point.
[0031] In a possible design of the first aspect, when the first item entity is the focus item entity, the electronic device marks the first text entity. The above method also includes: in response to a fourth trigger operation on the third item entity, adding a third text mark to the second text entity.
[0032] That is, as the focus item entity switches, the text entity marked by the electronic device will also switch, such as switching from the first text entity to the second text entity. In this way, the electronic device can always ensure that the marked text entity matches the current focus item entity.
[0033] In a possible design of the first aspect, the method further includes: in response to a move operation on the highlighted item entity, displaying a fourth interface, the fourth interface including multiple associated entries, each associated entry corresponding to an application or a service, the multiple associated entries including a first associated entry, the first associated entry corresponding to the first application or the second service. That is, for the focus item entity, the electronic device can quickly provide associated entries. In response to moving the highlighted item entity to the first associated entry, the electronic device displays a fifth interface, the fifth interface is an interface of the first application or the second service, and the fifth interface includes associated information of the highlighted item entity. That is, the user only needs to move the focus item entity to the first association, and the electronic device can present the associated information of the focus item entity in the first application or the second service.
[0034] In this way, the user's operation can be simplified. The user does not need to first exit the current user interface, such as the first interface, and enter the desktop, then enter the interface of the first application or the second service, and finally search for the focus item entity in the interface of the first application or the second service, thereby improving the efficiency of human-computer interaction.
[0035] In a possible design manner of the first aspect, the displaying of the fourth interface in response to the moving operation of the highlighted item entity includes: in response to the moving operation of the highlighted item entity, moving the position of the highlighted item entity in the first interface. In response to the position of the highlighted item entity moving to a target area in the first interface, displaying the fourth interface.
[0036] That is to say, the electronic device determines that there is a need to provide an associated entrance only after the focus object entity moves to the target area, thereby improving the accuracy of the timing of providing the associated entrance.
[0037] In a possible design of the first aspect, the highlighted item entity is a first item entity, and the multiple associated entries include a second associated entry. The highlighted item entity is a third item entity, and the multiple associated entries include a third associated entry. The second associated entry is different from the third associated entry.
[0038] That is to say, the associated entrances provided by the electronic device may be different depending on the focus object entity, so as to improve the specificity of the provided associated entrances.
[0039] In a possible design of the first aspect, the first interface is a camera viewfinder interface, and the first trigger operation includes a shooting operation. That is, the shooting operation can trigger the electronic device to recognize and mark the text entity in the interface.
[0040] In a possible design of the first aspect, before adding a first text tag to the first text entity in response to a first trigger operation on the first interface, the method further includes: displaying an identification control in the first interface when the first interface satisfies a second condition. The second condition includes: the first interface is a non-blank interface, the first interface includes text entities and / or object entities, and the first trigger operation includes a trigger operation on the identification control. That is, when the first interface includes useful information such as text and objects, the electronic device will actively push the identification control to identify and mark the text entity in the first interface. In this way, the electronic device can realize the function of intelligently pushing identification and marking text entities (such as the intelligent recognition function described below).
[0041] In a possible design of the first aspect, when the first interface satisfies the second condition, displaying the recognition control in the first interface includes: when the first interface satisfies the second condition, in response to a fifth trigger operation (such as a two-finger press operation) of the user on the first interface, displaying the recognition control in the first interface. The user performing the fifth operation on the first interface indicates that the user wants to recognize and mark the text entity.
[0042] That is to say, the electronic device can include useful information such as text and objects in the first interface, and only push the function of identifying and marking text entities when the user wants to identify and mark the text entities, so as to accurately meet the user's needs.
[0043] In a possible design manner of the first aspect, the method further includes:
[0044] In response to the fifth trigger operation of the user on the first interface, the fifth trigger operation is used to trigger the electronic device to analyze the content of the first interface to determine the functions required to be recommended for the first interface. Several typical contents and their matching functions are as follows:
[0045] When the number of foreign languages in the first interface exceeds the first number, a translation control is displayed in the first interface; when the number of foreign languages in the first interface does not exceed the first number, the translation control is not displayed, and the translation control is used to trigger the electronic device to translate the foreign languages in the first interface. In this way, the electronic device can recommend a translation function for interfaces with many foreign languages so as to translate the foreign languages in the interface.
[0046] When the first interface includes private information, a privacy protection control is displayed in the first interface; when the first interface does not include private information, the privacy protection control is not displayed, and the privacy protection control is used to trigger the electronic device to shield the private information in the first interface. In this way, the electronic device can recommend a privacy protection function for the interface including privacy to protect the private information in the interface.
[0047] When the first interface includes text but does not include an entity, a selection control is displayed in the first interface; when the first interface does not include text or includes an entity, the selection control is not displayed, and the selection control is used to trigger the electronic device to select the text in the first interface. In this way, the electronic device can recommend a text selection function for an interface including ordinary text, so as to perform operations such as copying and cutting on the text in the interface.
[0048] In all the above descriptions of the possible design methods of the first aspect, the first interface is mainly used for description. It can be understood that in various possible design methods, the specific implementation of the second aspect is similar to that of the first interface, and this article will not go into details.
[0049] In a second aspect, the present application further provides an electronic device, the electronic device comprising a display screen, a memory and one or more processors. The display screen, the memory and the processor are coupled. The memory is used to store computer program code, the computer program code comprising computer instructions, when the computer instructions are executed by the processor, the electronic device executes the method in the first aspect and any possible design thereof.
[0050] In a third aspect, the present application provides a chip system, which is applied to an electronic device including a display screen and a memory; the chip system includes one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected through lines; the interface circuit is used to receive signals from the memory of the electronic device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as in the first aspect and any possible design method thereof.
[0051] In a fourth aspect, the present application provides a computer storage medium, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method of the first aspect and any possible design thereof.
[0052] In a fifth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method of the first aspect and any possible design thereof.
[0053] It can be understood that the beneficial effects that can be achieved by the electronic device of the second aspect, the chip system of the third aspect, the computer storage medium of the fourth aspect, and the computer program product of the fifth aspect provided above can be referred to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A hardware structure diagram of an electronic device provided in an embodiment of the present application;
[0055] Figure 2 One of the mobile phone interface diagrams provided in the embodiment of the present application;
[0056] Figure 3 The second mobile phone interface diagram provided in the embodiment of the present application;
[0057] Figure 4 The third mobile phone interface diagram provided in the embodiment of the present application;
[0058] Figure 5 The fourth mobile phone interface diagram provided in the embodiment of the present application;
[0059] Figure 6The fifth mobile phone interface diagram provided in the embodiment of the present application;
[0060] Figure 7 The sixth mobile phone interface diagram provided in the embodiment of the present application;
[0061] Figure 8 The seventh mobile phone interface diagram provided in the embodiment of the present application;
[0062] Fig. 9A The eighth mobile phone interface diagram provided in the embodiment of the present application;
[0063] Fig. 9B Figure 9 of the mobile phone interface provided in the embodiment of the present application;
[0064] Fig.10 The tenth mobile phone interface diagram provided in the embodiment of the present application;
[0065] Fig.11 The eleventh mobile phone interface diagram provided in the embodiment of the present application;
[0066] Fig.12 The twelfth mobile phone interface diagram provided in the embodiment of the present application;
[0067] Figure 13A-13B The thirteenth mobile phone interface diagram provided for the embodiment of the present application;
[0068] Fig.14 The fourteenth mobile phone interface diagram provided for the embodiment of the present application;
[0069] Fig.15 The fifteenth mobile phone interface diagram provided in the embodiment of the present application;
[0070] Fig.16 Figure 16 of the mobile phone interface provided in the embodiment of the present application;
[0071] Fig.17 Figure 17 of the mobile phone interface provided in the embodiment of the present application;
[0072] Fig.18 This is the eighteenth mobile phone interface diagram provided for the embodiment of the present application. DETAILED DESCRIPTION
[0073] The technical solutions in the embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to be used as limitations on the present application. As used in the specification and the appended claims of the present application, the singular expressions "a", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" refer to one or more (including two). The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in a "or" relationship.
[0074] References to "one embodiment" or "some embodiments" etc. described in this specification mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Thus, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise specified. "First" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.
[0075] In the embodiments of the present application, the words "exemplarily" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.
[0076] The screen recognition method provided in the embodiment of the present application enables the electronic device to recognize the content in the user interface, such as text, objects, graphic codes (including bar codes, QR codes, etc.), and mark part of the content in the user interface based on the recognition results, thereby providing users with marks for quickly obtaining information from the user interface, which is conducive to improving the efficiency of human-computer interaction.
[0077] Taking a picture displayed in the user interface as an example, the electronic device can recognize the text, objects and graphic codes included in the picture, and based on the recognized objects, mark some of the text and graphic codes in the picture, such as marking the text and graphic codes that are highly correlated with the objects.
[0078] For example, the electronic device may be a mobile phone, tablet, desktop, laptop, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) and virtual reality (VR) device, etc., which have a camera. The embodiments of the present application do not impose any special restrictions on the specific form of the electronic device.
[0079] See also Figure 1 The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0080] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0081] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0082] In some embodiments, the electronic device can complete the recognition method through the processor 110 to obtain a recognition result.
[0083] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.
[0084] The electronic device implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.
[0085] The display screen 194 is used to display images, videos, etc. In some embodiments, the electronic device can display the recognition result through the display screen 194.
[0086] The electronic device can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194 and the application processor, etc. In some embodiments, the electronic device can collect images through the camera 193 for recognition.
[0087] The electronic device can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0088] The button 190 may include a power button, a volume button, etc. The button 190 may be a mechanical button. It may also be a touch button. The mobile phone may receive the button input and generate a key signal input related to the user settings and function control of the mobile phone. The motor 191 may generate a vibration prompt. The motor 191 may be used for incoming call vibration prompts or for touch vibration feedback. The indicator 192 may be an indicator light, which may be used to indicate the charging status, power changes, messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card. The SIM card may be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the mobile phone.
[0089] The software system of the electronic device may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture, which is not specifically limited in the embodiments of the present application.
[0090] The screen recognition method provided in the embodiment of the present application can be implemented in the above-mentioned electronic device. The following takes the electronic device being a mobile phone as an example to illustrate the screen recognition method provided in the embodiment of the present application.
[0091] The mobile phone provides an application (Application, APP) specifically used to identify the content in the user interface, which is referred to as the smart vision APP in this article. The smart vision APP can be a system-level APP. Among them, the user interface includes the application interface of each application, such as the lock screen interface, desktop, gallery interface, chat interface, video playback interface, etc.
[0092] In particular, the application interface also includes the camera application's framing interface. In other words, the Smart Vision APP can also be used to identify framing content.
[0093] There are many ways to open the Smart Vision APP on your mobile phone, as shown in Method 1 to Method 4 below.
[0094] Method 1: The mobile phone interface provides an opening entrance for the Smart Vision APP.
[0095] The mobile phone provides an opening entry for the Smart Vision APP. In response to a trigger operation on the opening entry, such as a click operation, the mobile phone can open the Smart Vision APP for identification. The embodiment of the present application does not specifically limit the form and position of the opening entry. This article only lists the following Figure 2 Several example positions are shown:
[0096] Example 1: The opening entrance is located on the lock screen interface.
[0097] In the lock screen interface of the mobile phone, in response to the user pulling up from the bottom of the lock screen interface, the mobile phone can display the entrances to various functions provided by the mobile phone in the lock screen interface. For example, in response to the pull-up operation, the mobile phone can display Figure 2 The lock screen interface 201 shown. The lock screen interface 201 includes the entrances of the quick functions such as the recorder, calculator, and compass. In addition, the lock screen interface 201 also includes the opening entrance 200 of the smart vision APP.
[0098] Example 2: The opening entrance is located in the drop-down control center.
[0099] When the phone screen is on, in response to the user pulling down from the top of the display screen, the phone can display Figure 2 The pull-down control center interface 202 shown. The pull-down control center interface 202 also includes entrances to various shortcut functions provided by the mobile phone, including an opening entrance 200 of the Smart Vision APP.
[0100] Example 3: The opening entrance is located in the global search interface.
[0101] When the mobile phone displays the desktop (including the main screen and the negative one screen), in response to the user sliding down from the middle area of the desktop, the mobile phone can display Figure 2 The global search interface 203 shown. The global search interface 203 can be used to search for information in the mobile phone. For example, the global search interface 203 includes a search box 2031. In response to the user inputting a search text in the search box, the mobile phone can search for information including the searched text in the mobile phone, such as chats, files, text messages, applications, etc. The search box 2031 includes an opening entry 200 of the Smart Vision APP.
[0102] Example 4: the opening entrance is located at the negative one screen.
[0103] The phone can display Figure 2 The negative one screen 204 shown includes a search box 2041 , and the search box 2041 includes an opening entrance 200 of the Smart Vision APP.
[0104] Example 5: The opening entrance is located in the application interface of the camera application.
[0105] After opening the camera app on your phone, the phone can display Figure 2 The application interface 205 of the camera application is shown. The application interface 205 includes an opening entry 200 of the smart vision APP.
[0106] Method 2: In response to voice wake-up, the mobile phone opens the Smart Vision APP.
[0107] The user can also trigger the mobile phone to open the smart vision APP by voice wake-up. Specifically, after the voice assistant of the mobile phone is awakened, in response to the user inputting the voice to open the smart vision APP, the mobile phone can open the smart vision APP for recognition.
[0108] Take the voice input of "Use Smart Vision" as an example. After the user inputs the voice input "Use Smart Vision", the phone can recognize "Use Smart Vision" and display Figure 3 The interface 301 shown includes a voice result "Use Smart Vision" 3011. Subsequently, the mobile phone can open the Smart Vision APP.
[0109] After the Smart Vision APP is triggered and opened through the above method 1 or method 2, the mobile phone can display the application interface of the Smart Vision APP. For example, the mobile phone can display Figure 4 The interface 401 shown is the application interface of the smart vision APP. The interface 401 includes a viewfinder area 4010 for displaying the viewfinder image of the camera.
[0110] Furthermore, the Smart Vision APP provides a variety of recognition functions. For example, Figure 4 The interface 401 shown includes multiple recognition function options, such as "text extraction" option 4011, "intelligent recognition" option 4012, "scan code" option 4013 (usually the default selected option), and "cutout" option 4014. In response to the user's sliding operation from right to left in the function option area 4015 (shown by the dotted line box in the figure, which does not actually exist), the interface 401 may also present a "translation" option 4016, a "scan file" option 4017, a "scan card" option 4018, and a "test paper / homework" option 4019.
[0111] In response to the user's selection operation of any function option, the electronic device can switch to the corresponding recognition function to achieve recognition. Figure 4 By clicking the "Smart Identification" option 4012 in the interface 401, the mobile phone can switch from the default scanning function to the smart identification function, such as displaying Figure 4 In the interface 402 shown, the currently selected item in the interface 402 is “intelligent identification” 4012 .
[0112] The text extraction function is used to identify and extract text in the user interface, and can also be used to mark text entities. Text entities include: ID numbers for identification (such as ID card numbers, passport numbers, etc.), addresses, phone numbers, flight information, express numbers, email addresses, links, etc. For example, when a phone number is identified in an image, a phone mark is added to the phone number.
[0113] Intelligent recognition function: On the one hand, the intelligent recognition function can integrate the above-mentioned text extraction function; on the other hand, the intelligent recognition function can also be used to identify objects in the user interface and mark the object entities. Among them, object entities include: animals, plants, buildings, food, etc. For example, when an animal is identified in the picture, an animal mark is added to the corresponding area. In other words, the intelligent recognition function can realize the recognition and marking of text, as well as the recognition, cutting and labeling of objects.
[0114] In addition, the intelligent recognition function can also recognize and mark graphic codes, including barcodes, QR codes, etc.
[0115] The translation feature can be used to recognize text in the user interface and then translate it.
[0116] After triggering the opening of the Smart Vision APP through the above-mentioned method 1 or method 2, the objects identified by the corresponding recognition function include: pictures taken under the corresponding recognition function, or pictures selected from the gallery, which is not specifically limited in the embodiments of the present application.
[0117] Exemplarily, in response to a user's Figure 4 By clicking the shooting button 4020 in the interface 402 (also referred to as a first trigger operation or a second trigger operation), the mobile phone can take a picture and use the currently selected recognition function (ie, the intelligent recognition function) to recognize the taken picture to obtain a recognition result.
[0118] In another exemplary embodiment, in response to a user clicking on the gallery entry 4021 in the interface 402, the mobile phone may provide pictures in the gallery for the user to select, such as displaying Figure 4 Interface 403 shown includes thumbnails of multiple pictures in the gallery. In response to a user's selection operation on any thumbnail, such as thumbnail 4031, the mobile phone can use the currently selected recognition function to recognize the picture corresponding to the selected thumbnail to obtain a recognition result.
[0119] And, after the Smart Vision APP is triggered to start through the above method 1 or method 2, the objects identified by the corresponding recognition function also include: the content that the mobile phone has framed but not yet photographed under the corresponding recognition function, that is, the content of the framing interface. For example, the identified object can be Figure 4 The content corresponding to the framing screen 4022 displayed in the framing interface 402 is shown. In this case, the mobile phone can recognize the content of the framing interface through augmented reality (AR) recognition technology to obtain a recognition result.
[0120] Method 3: After launching the Gallery app or displaying a picture in the Gallery app, the phone automatically launches the Smart Vision app.
[0121] The mobile phone can also automatically start the Smart Vision APP to analyze the pictures in the Gallery App after starting the Gallery App or displaying the pictures in the Gallery App, and recommend the intelligent recognition function in the Smart Vision APP based on the analysis results.
[0122] In some embodiments, after starting the gallery application, the mobile phone can start the Smart Vision APP, use the Smart Vision APP to parse the pictures in the gallery, and determine the target picture. Among them, the target picture can include at least one of the following pictures: a non-blank picture, a picture with an object, a picture with text, and a picture with a graphic code. In other words, the target picture usually contains useful information such as text, objects, and graphic codes. Pictures that are not target pictures do not include the above useful information.
[0123] In response to a user's viewing operation on a target image, such as a click operation on a thumbnail corresponding to the target image, the mobile phone can display the target image and provide an entry to the smart recognition function. The entry to the smart recognition function can be used to trigger the mobile phone to recognize and mark the text, object, graphic code, etc. in the target image. In this way, in response to the viewing operation, the mobile phone can quickly recommend the smart recognition function and provide the user with an entry to recognize and mark the content in the image.
[0124] For example, after starting the gallery application, the phone can display Figure 5 Interface 501 is an application interface of a gallery application. Interface 501 includes thumbnails of multiple pictures in the gallery, such as thumbnail 5011. Taking the target picture as the picture corresponding to thumbnail 5011 as an example, in response to the user clicking on thumbnail 5011, the mobile phone can display Figure 5 The interface 502 shown in FIG. The interface 502 includes a picture 5021 corresponding to the thumbnail 5011 and also includes "Smart Vision" 5022, which is the entrance to the smart recognition function. That is, the mobile phone recognizes that the picture 5021 is the target picture.
[0125] After each launch of the gallery, the phone can only analyze the newly added pictures in the phone between the two launches of the gallery, and other pictures can refer to the historical recognition results. In this way, after each launch of the gallery, the phone can only analyze a small number of pictures, thereby reducing the power consumption of the phone.
[0126] In other embodiments, after receiving a user's operation of viewing a picture in the gallery application, the mobile phone can start the Smart Vision APP and use the Smart Vision APP to analyze the currently viewed picture to determine whether it is a target picture. In this way, the mobile phone can only analyze the currently viewed picture, thereby reducing power consumption for each analysis.
[0127] Of course, the timings for triggering the mobile phone to automatically start the Smart Vision APP in the above method three are only a few typical situations, and are not actually limited to this. For example, the mobile phone can also automatically start the Smart Vision APP to analyze the pictures in the gallery when charging or at a preset time (such as early morning) to determine the target icon. On this basis, after starting the gallery application, the mobile phone can start the Smart Vision APP to analyze pictures that have not been analyzed, that is, the pictures newly added between the last analysis and the current startup of the gallery application; or, on this basis, the mobile phone only starts the Smart Vision APP to analyze the currently viewed picture after receiving the user's viewing operation on the picture that has not been analyzed in the gallery application.
[0128] If the currently viewed picture is the target picture, the phone can provide an entry to the smart recognition function. If the currently viewed picture is not the target picture, the phone does not provide an entry to the smart recognition function. Figure 5 The thumbnail 5012 in the interface 501 shown is taken as the second thumbnail. In response to a click operation on the thumbnail 5012, the mobile phone may display Figure 5 The interface 504 shown in FIG. 504 includes a picture 5041 corresponding to the thumbnail 5012. However, the interface 504 does not include an entry for the intelligent recognition function. That is, the interface 504 is the fourth interface.
[0129] After automatically starting the Smart Vision APP and providing the entrance to the smart recognition function in the third method, if the display time of the entrance to the smart recognition function reaches time 1, the mobile phone can retract the entrance to the smart recognition function to the edge of the user interface to avoid affecting the user's viewing of pictures. Figure 5 After the display time of "Smart Vision" 5022 in the interface 502 reaches 5 seconds, the mobile phone can display Figure 5 Interface 503 is shown. Interface 503 also includes "Smart Vision" 5022. However, unlike interface 502, "Smart Vision" 5022 in interface 503 is displayed at the edge of interface 503. It should be noted that the function of the closed smart recognition function entrance is the same as that of the expanded smart recognition function entrance, and both can trigger the mobile phone to use the smart recognition function to recognize the target image, which will not be repeated here.
[0130] Furthermore, during a use process from starting the gallery process to closing the gallery process, after receiving the user's trigger operation on the entrance of the recommended smart recognition function, even if the mobile phone responds to the user's viewing operation on the target picture again, the mobile phone may no longer display the expanded smart recognition function entrance, but directly display the collapsed smart recognition function entrance.
[0131] After automatically starting the Smart Vision APP and providing an entry to the smart recognition function using the third method, in response to a user's triggering operation (which may be referred to as a first triggering operation or a second triggering operation) on the entry to the smart recognition function, such as a click operation, the mobile phone may use the smart recognition function to recognize the target image and obtain a recognition result. Figure 5 By clicking on "Smart Vision" 5022 in the interface 502 shown, the mobile phone can use the intelligent recognition function provided by the Smart Vision APP to recognize the image 5021 and obtain the recognition result.
[0132] Method 4: In response to the user's operation 1 (also referred to as the fifth trigger operation), the mobile phone starts the Smart Vision APP.
[0133] After receiving the user's operation 1 in the current interface (which can be called the first interface or the second interface), the mobile phone can also start the smart vision APP to analyze the content of the current interface and recommend processing functions based on the analysis results. That is to say, in method 4, the mobile phone can be triggered by operation 1 to start the smart vision APP to analyze the corresponding needs and recommend.
[0134] Among them, operation 1 can be a long press operation, a sliding operation, a two-finger pressing operation, etc. This embodiment of the present application does not specifically limit this. The following mainly takes operation 1 being a two-finger pressing operation as an example for explanation.
[0135] The current interface may be any user interface displayed during the use of the mobile phone. For example, the current interface may be an application interface of a social application (such as a chat application, a life sharing application, etc.), a picture viewing interface, etc. The present application embodiment does not specifically limit this. That is to say, in method 4, the mobile phone can not only intelligently recommend processing functions for pictures in the gallery application, but also recommend processing functions for other interfaces.
[0136] At this point, it should be noted that: the current interface is a picture viewing interface, and the mobile phone can directly use the picture as the content of the current interface for the smart vision APP to analyze. However, if the current interface is not a picture viewing interface, it is usually difficult for the mobile phone to directly obtain the content of the current interface. Based on this, in a specific implementation method, in response to the user's two-finger press operation on the current interface, the mobile phone can take a screenshot of the current interface and start the smart vision APP to analyze the screenshot. That is, the mobile phone can obtain the content of the current interface by taking a screenshot and the user's smart vision APP analyzes it. Correspondingly, the smart vision APP analyzes the screenshot, which is equivalent to analyzing the content of the current interface.
[0137] In another specific implementation, in response to the user's two-finger pressing operation on the current interface, and when the current interface is a picture viewing interface, the mobile phone can start the Smart Vision APP to analyze the currently viewed picture. In response to the user's two-finger pressing operation on the current interface, and when the current interface is not a picture viewing interface, the mobile phone can take a screenshot of the current interface and start the Smart Vision APP to analyze the screenshot. In this way, the mobile phone can take targeted screenshots for analysis by the Smart Vision APP.
[0138] In the following, the example in which a mobile phone takes a screenshot of the current interface and starts the Smart Vision APP to analyze the screenshot in response to a two-finger press operation on the current interface is mainly used for explanation. The following content of this application can also be a technical solution after parsing the image or interface content after starting the Smart Vision APP using Methods 1 to 3.
[0139] The processing function may be a recognition function provided by the smart vision APP, such as a translation function, an intelligent recognition function, or a text extraction function. Alternatively, the processing function may be other functions, such as a privacy protection function. The privacy protection function is used to identify and block private information to protect it.
[0140] In some embodiments, after the number of foreign languages in the current interface is parsed by the Smart Vision APP to reach a threshold number (also referred to as a first number), the mobile phone may provide an entry for a translation function (also referred to as a translation control). In this way, the mobile phone may recommend the translation function in the Smart Vision APP for scenarios where translation is required.
[0141] The quantity threshold may be a fixed quantity, or the quantity of text included in the current interface (i.e., the interface is all in foreign languages), or a fixed ratio (such as 90%, 80%, etc.) of the quantity of text included in the current interface. This embodiment of the application does not specifically limit this.
[0142] Taking the quantity threshold as the number of texts included in the current interface as an example, the mobile phone can display Figure 6 In the interface 601 shown, the interface 601 is all in English. In response to the user's two-finger press operation on the interface 601, the mobile phone can take a screenshot of the interface 601 and start the Smart Vision APP to analyze the screenshot. After parsing that the text included in the screenshot is all in foreign language (different from the language set by the system), the mobile phone can display Figure 6 The interface 602 shown is different from the interface 601 in that the interface 602 includes a "Translation" 6021. The "Translation" 6021 is the entry of the translation function.
[0143] On the contrary, if the Smart Vision APP analyzes that the number of foreign languages in the current interface does not reach the threshold, the phone will not provide an entry for the translation function. In other words, the phone will not recommend the translation function for each interface when you press two fingers, but will recommend it dynamically based on the number of foreign languages.
[0144] In other embodiments, after the Smart Vision APP analyzes the current interface to include privacy information, the mobile phone can provide an entry to the privacy protection function (also called a privacy protection control). In this way, the mobile phone can recommend a privacy protection function for scenes with privacy protection requirements.
[0145] Among them, private information includes address, telephone number, email address, ID number / picture for identity identification, etc.
[0146] For example, the phone may display Figure 7 Interface 701 is an application interface of a chat application. In response to the user's two-finger pressing operation on interface 701, the mobile phone can take a screenshot of interface 701 and start the Smart Vision APP to analyze the screenshot. After analyzing the screenshot to find that the screenshot includes private information such as address and email address, the mobile phone can display Figure 7 The interface 702 shown is different from the interface 701 in that the interface 702 includes a "privacy code" 7021. The "privacy code" 7021 is the entrance to the privacy protection function.
[0147] On the contrary, if the Smart Vision APP analyzes that the current interface does not include privacy information, the phone will not provide an entry to the privacy protection function. In other words, the phone will not recommend the privacy protection function for each interface when two-finger press is performed, but will dynamically recommend it based on whether the current interface includes privacy information.
[0148] In other embodiments, after the Smart Vision APP is used to parse that the current interface includes text but does not include entities (including text entities and object entities), the mobile phone can provide an entry for the selection function (also called a selection control). It should be noted that the mobile phone can select text only after the text is extracted, so the selection function can be understood as a sub-function of the aforementioned text extraction function.
[0149] It is understandable that if there is no entity in the current interface, it means that the user has no need to perform operations on the entity, and the mobile phone can directly provide an entry for the selection function. In this way, the mobile phone can quickly perform a selection operation on the text in the current interface, so as to further copy, cut or share the text later.
[0150] For example, the phone may display Figure 8Interface 801 shown in FIG. 8 includes picture 8011. Picture 8011 includes text but no entity. In response to the user's two-finger press operation on interface 801, the mobile phone can take a screenshot of interface 801 and start the Smart Vision APP to analyze the screenshot. After analyzing that the screenshot includes text but no entity, the mobile phone can display Figure 8 The interface 802 shown is different from the interface 801 in that the interface 802 includes a "select all" 8021. The "select all" 8021 is the entry for selecting a function.
[0151] On the contrary, if the Smart Vision APP analyzes that the current interface does not contain text, or contains text and entities, the phone will not provide an entry for the selection function. In other words, the phone does not recommend the selection function for each interface when two-finger press is performed, but will recommend dynamically based on whether the current interface contains text and / or entities.
[0152] The above-mentioned solution of the recommendation selection function is particularly suitable for some interfaces where the copy operation cannot be performed, such as the above-mentioned interface 801, or some application interfaces that restrict users from copying, so that users can complete the operation on the text in these interfaces.
[0153] It can be seen that, in the fourth method, for different interfaces, in response to the user's two-finger pressing operation on the interface, the mobile phone can use the smart vision APP to parse the user interface and recommend different recognition functions. For example, the current interface is interface 2, in response to the user's two-finger pressing operation on interface 2, the mobile phone can recommend processing function 1; the current interface is interface 3, in response to the user's two-finger pressing operation on interface 3, the mobile phone can recommend processing function 2. Among them, processing function 1 and processing function 2 are different processing functions.
[0154] Furthermore, the same interface may satisfy the recommendation conditions of different processing functions, such as satisfying the recommendation conditions of the translation function and the selection function at the same time. In this case, the mobile phone can recommend all the processing functions that meet the conditions at the same time; or, the mobile phone can configure the matching order of different (types) of interfaces and processing functions, and the mobile phone can only recommend one processing function that has the highest matching degree with the current interface. The embodiments of the present application do not specifically limit this.
[0155] For example, for a user interface that does not allow long press to copy, the mobile phone can be configured with a selection function with the highest match degree to facilitate text operations on the user interface; for an application interface of a chat application, the mobile phone can be configured with a privacy protection function with the highest match degree to facilitate coding of private information involved in the chat content before sending it; and, for foreign language websites, the mobile phone can be configured with a translation function with the highest match degree to facilitate foreign language translation.
[0156] After automatically starting the Smart Vision APP and providing an entrance to the corresponding processing function using the aforementioned method 4, in response to the user's trigger operation on the entrance to the processing function (which can be called a first trigger operation or a second trigger operation), such as a click operation, the mobile phone can use the recommended processing function to process the current interface.
[0157] At this point, it should be noted that the conditions for triggering the launch of the Smart Vision APP in the above-mentioned method three and method four can also be interchanged. That is, launching the gallery or receiving an operation to view pictures in the gallery in method three can be interchanged with the double-finger pressing operation in method four. Exemplarily, in response to launching the gallery or receiving an operation to view pictures in the gallery, the mobile phone can launch the Smart Vision APP to analyze whether the conditions for recommending various processing functions in the above-mentioned method four are met. If so, the corresponding processing functions are recommended. Another exemplary example is that in response to a double-finger pressing operation, the mobile phone can launch the Smart Vision APP and analyze whether the conditions for recommending the intelligent recognition function in the above-mentioned method three are met. If so, the intelligent recognition function is recommended.
[0158] After going through the aforementioned methods one to four, the mobile phone can use corresponding functions (such as various recognition functions provided by the smart vision APP selected in method one and method two, or the intelligent recognition function provided by the smart vision APP recommended in method three, and various processing functions recommended in method four) to perform processing.
[0159] It should be noted that when the mobile phone uses the corresponding function to perform processing, it needs to identify the text, objects, graphic codes, etc. in the current interface, and mark, select, translate, and other operations, which will change the information in the current interface. Based on this, in a specific implementation method, in response to the user triggering the operation of using the corresponding function to perform processing, such as the click operation on the shooting button 4020 in method one and method two, the selection operation on the thumbnail 4031, the click operation on the "Smart Vision" 5022 in method three, and the trigger operation on the entrance of various recommended processing functions in method four, the mobile phone can first screenshot the current interface and display the screenshot image, and identify the screenshot image, and then perform various operations such as marking, selecting, translating, etc. In this way, the mobile phone can perform operations on the screenshot image without affecting the current interface. It should be noted that in the process of browsing large images in the gallery and identifying and processing the image content, the mobile phone may not perform screenshot processing.
[0160] The following describes the intelligent recognition function, text extraction function, translation function and privacy protection function respectively. It should be noted that in the following, in response to the user triggering the operation of using the corresponding function to perform the processing, the process of the mobile phone taking a screenshot and performing the operation on the screenshot image will not be described one by one.
[0161] First, intelligent recognition function.
[0162] The mobile phone uses intelligent recognition function to identify and mark object entities.
[0163] Taking the above method as an example, after selecting the function option of the smart recognition function, the mobile phone can display Fig. 9A In response to the user clicking the capture button 9011 in the interface 901, the mobile phone may display Fig. 9A Interface 902 is shown, and interface 902 is the recognition result page (also referred to as the first interface). In the embodiment of the present application, in response to the user clicking the shooting button in interface 901 corresponding to the intelligent recognition function, the mobile phone can take a photo, but does not save the result of the photo. The content displayed in interface 902 is the cached photo result. If the user clicks the return control in interface 902 and does not save the result of the photo, then the mobile phone will not store the result of the photo; or, if the user continues to edit the content of interface 902 and chooses to save, the mobile phone can also save the corresponding result. Various object entities are marked in interface 902 using circles, animal icons and other symbols. For example, dog 9021 is marked with animal icon 90211, building 9022 is marked with circle 90221, and plant 9023 is marked with circle 90231.
[0164] In the recognition results, the mobile phone can highlight the item entity with the largest area and / or the most accurate recognition as the focus item entity (herein, the highlight effect is represented by a bold line, but it is not limited to this), and mark the focus item entity with a category icon. For example, the recognition result of dog 9021 in interface 902 is the most accurate, so dog 9021 is highlighted in interface 902; and the upper left corner of dog 9021 is marked with an animal icon 90211, which clearly indicates that the current focus item entity is an animal, that is, animal icon 90211 is a category icon.
[0165] Other object entities in the recognition result are not highlighted, and are marked with common symbols. For example, the building 9022 and the plant 9023 in the interface 902 are not highlighted, and are marked with circles, with the circle 90221 marking the building 9022 and the circle 90231 marking the plant 9023.
[0166] After the recognition result is displayed, in response to the operation of switching the focus (which may be referred to as the fourth trigger operation), the mobile phone may switch the focus item entity. That is, the focus is switched from the current focus item entity (which may be referred to as the first item entity) to another item entity (which may be referred to as the third item entity). Similarly, the switched focus item entity is highlighted, and the switched focus item entity is marked with a category icon.
[0167] In some embodiments, the operation of switching focus may be a triggering operation, such as a click operation, on a mark corresponding to a target item entity other than the current focus item entity (such as a first item mark of a third item entity).
[0168] The current focused item entity is Fig. 9A The dog 9021 in the interface 902 shown, the target object entity is Fig. 9A For example, in the interface 902 shown in FIG. 902 , in response to the user clicking on the circle 90221 in the interface 902 , the mobile phone may display Fig. 9A Interface 903 is shown. The difference from interface 902 is that in interface 903, building 9022 is highlighted, and a building icon 9031 is marked in the upper left corner of building 9022, which clearly indicates that the focus item entity after switching is the building, that is, building icon 9031 is a category icon; at the same time, in interface 903, dog 9021 is no longer highlighted, and a circle is also used, such as circle 9032 to mark dog 9021. That is, the focus item entity is switched from dog 9021 to building 9022.
[0169] In other embodiments, the operation of switching focus includes a click operation on a mark corresponding to a target item entity other than the current focus item entity, and a click operation on the current focus item entity.
[0170] The current focused item entity is Fig. 9B The dog 9021 in the interface 902 shown, the target object entity is Fig. 9B For example, in the interface 902 shown in FIG. 902 , in response to the user clicking on the circle 90221 in the interface 902 , the mobile phone may display Fig. 9B The interface 905 shown in FIG. 9 is different from the interface 902 in that the interface 905 not only highlights the dog 9021, but also the building 9022. Subsequently, in response to the user's click operation on the dog 9021 in the interface 905, the mobile phone may display Fig. 9B Interface 903 is shown. Different from interface 905, dog 9021 is no longer highlighted in interface 903, but only building 9022 is highlighted. In this way, the focus can be switched from dog 9021 to building 9022.
[0171] Further, in this embodiment, in response to the user's Fig. 9B By clicking on the building 9022 in the interface 903, the mobile phone can display Fig. 9B The interface 906 shown further cancels the highlighting of the building 9022. That is, in response to the user's click operation on the focus item entity, the highlighting of the focus item entity can be canceled, that is, the focus item entity is switched to non-focus.
[0172] In response to a user triggering operation on a category icon of a focus item entity, such as a click operation, the mobile phone may display introduction information of the focus item entity to assist the user in understanding the focus item entity. Fig. 9A By clicking the building icon 9031 in the interface 903, the mobile phone can display Fig. 9A Interface 904 is shown. Different from interface 903 , interface 904 includes a pop-up window 9041 , and pop-up window 9041 includes introduction information of building 9022 .
[0173] The mobile phone can also display a quick entry around the focus item entity to implement quick operation on the focus item entity. For ease of description, the quick entry of the first item entity can be called the first quick entry, and the quick entry of the third item entity can be called the second quick entry.
[0174] Among them, the quick entrance can be fixed, such as a search entrance, a purchase entrance, a copy entrance, a save entrance, a share entrance, etc.
[0175] Alternatively, the quick entry can be different for different focus item entities. For example, for buildings, the mobile phone can provide a search entry; for commodities, the mobile phone can provide a purchase entry and a price comparison entry. In this way, the mobile phone can provide targeted quick entry to accurately meet the needs of users.
[0176] The focused item entity is Fig. 9A Taking the building 9022 in the interface 903 as an example, the following quick access entries are displayed at the lower edge of the building 9022: “Search” 9033, “Save” 9034 and “Share” 9035.
[0177] In response to the user's triggering operation on the quick entry, such as a click operation, the mobile phone can complete the corresponding quick operation for the focus item entity. Among them, in response to the user's click operation on the search entry, the mobile phone can search for information related to the focus item entity in the network and display it. Exemplarily, in response to the user's click operation on "Search" 9033 in interface 903, the mobile phone can search for relevant information about building 9022. For example, after the search is completed, the mobile phone can display Fig. 9A Interface 904 is shown. The introduction information in the pop-up window 9041 of interface 904 is obtained by searching.
[0178] Further, in response to the user's moving operation on the focus item entity, such as a long press followed by a drag operation, the mobile phone can move the display position of the focus item entity. It should be noted that after moving the display position of the focus item entity, the mobile phone can still display the focus item entity at the initial position of the focus item entity. Of course, the mobile phone may also not display the focus item entity at the initial position, and the embodiments of the present application do not specifically limit this.
[0179] The focused item entity is Fig. 9A Taking the dog 9021 in the interface 902 as an example, in response to the user's long press and drag operation on the dog 9021, the mobile phone can display Fig.10 The interface 1001 shown is different from the interface 902 in that the position of the dog 9021 in the interface 1001 has changed.
[0180] In response to the display position of the focus item entity after moving to the target area in the user interface, the mobile phone displays multiple shortcut icons (also called associated entrances) of associated applications / functions / services, so as to quickly execute the associated application / function-related processing for the focus item entity. For the sake of convenience, the interface displaying the shortcut icon can be called the fourth interface. The following mainly takes the associated application as an example, and the corresponding shortcut icon is the application icon.
[0181] The target area may be an edge area of the user interface, such as a left edge or a right edge.
[0182] Among them, the associated applications can be fixed, such as chat applications, search applications, shopping applications, price comparison applications, sharing applications, collection applications, printing applications, recipe applications, cute pet applications, etc. Alternatively, corresponding to different focus item entities, the associated applications may not be exactly the same, that is, they may include different shortcut icons. Exemplarily, for commodities, the associated applications may include purchase applications and price comparison applications; for food, the associated applications may include recipe applications; for animals, the associated applications may include cute pet applications. The associated functions or services can also be displayed in a similar manner, which will not be repeated here.
[0183] Continue to see Fig.10 As the user continues to move the dog 9021 in the interface 1001, the dog 9021 may be moved to Fig.10 In the area 10021 (indicated by a dotted box in the figure, which does not exist in reality) in the interface 1002 shown, the area 10021 is the target area. The interface 1002 also includes shortcut icons 10022 for favorite applications, 10023 for search applications, 10024 for cute pet applications, 10025 for sharing services, 10026 for chat applications, and other related application shortcut icons.
[0184] In a specific implementation, the mobile phone can also display a dynamic effect when moving the display position of the focus object entity. Fig.10 The picture in interface 1001 occupies the entire user interface, and the picture in interface 1002 forms a "door" animation.
[0185] In a specific implementation, after the focus object entity is moved into the target area, the mobile phone can shrink the focus object entity. For example, the dog 9021 in the interface 1002 is much smaller than the dog 9021 in the interface 1001.
[0186] After displaying the shortcut icon of the associated application, in response to the focus item entity being moved to the position of the target shortcut icon (also referred to as the first associated entry), the mobile phone can display interface 1 (also referred to as the fifth interface) of the target associated application (also referred to as the first application). The target associated application is an associated application among multiple associated applications that corresponds to the target shortcut icon, and interface 1 includes information about the focus item entity.
[0187] The target shortcut icon is Fig.10 As an example, the shortcut icon 10024 in the interface 1002 shown in the figure is used, that is, the target associated application is the cute pet application, and the interface 1 is the application icon of the cute pet application. As the user continues to move the dog 9021 in the interface 1001, the mobile phone can display Fig.10 In the interface 1003 shown, the dog 9021 in the interface 1003 is moved to the position of the shortcut icon 10024. In response to the dog 9021 being moved to the position of the shortcut icon 10024, the mobile phone can display Fig.10 The interface 1004 shown includes a floating window 10041 , in which the application interface of the cute pet application is displayed, and the application interface of the cute pet application includes information related to the dog 9021 .
[0188] At this point, it should be noted that: in the above specific implementation method of recommending shortcut icons, the user needs to drag the focus item entity to the target area before triggering the display of the shortcut icon. In practice, it is not limited to this. Exemplarily, in response to the user's movement operation of the focus item entity, and the movement distance exceeds the distance threshold, the mobile phone can display the shortcut icon. Another exemplary embodiment, in response to the user's movement operation of the focus item entity, the mobile phone can display the shortcut icon.
[0189] The above introduction to the intelligent recognition function mainly describes the recognition of object entities, such as dog 9021, building 9022, and plant 9023, and the subsequent processing based on the recognition results. Based on the previous introduction to the intelligent recognition function, it can be seen that the intelligent recognition function can also be used for text recognition. For the relevant characteristics of the intelligent recognition function for text recognition, please refer to the following introduction to the text extraction function. I will not explain it in detail here.
[0190] Second, text extraction function.
[0191] The phone uses a text extraction function that can recognize and extract text and can also mark text entities.
[0192] Taking the above method 1 as an example, after selecting the function option of the text extraction function, the mobile phone can display Fig.11 In response to the user clicking the shooting button 11011 in the interface 1101, the mobile phone may display Fig.11 The interface 1102 shown is a recognition result page. In the interface 1102, address 1 is marked with a location icon 11021, website 1 is marked with a network icon 11022, and phone 2 is marked with a phone icon 11023.
[0193] In the recognition result, the mobile phone can mark all text entities. Alternatively, the mobile phone can only mark some text entities to avoid confusion.
[0194] In a specific implementation, for text entities of the same category, the mobile phone may mark only the first occurrence of the text of the category, wherein the mobile phone marks in the order from top to bottom and from left to right of the interface, and the first occurrence refers to the first occurrence from top to bottom and from left to right.
[0195] For example, the mobile phone executes the text extraction function and after recognizing the current interface, it can display Fig.12 Interface 1201 is shown. Interface 1201 is a recognition result page. In interface 1201, the first appearing address, i.e., Address 1, is marked with a location icon 12011, the first appearing phone number, i.e., Phone 1, is marked with a phone icon 12012, and the first appearing website link, i.e., Website 1, is marked with a network icon 12013.
[0196] That is, if the entity categories of two text entities (which may be referred to as the first text entity and the fifth text entity) are the same, only one of the text entities (such as the first text entity) may be marked. If the entity categories of the two text entities are different, the two text entities may be marked separately. For ease of explanation, the mark of the first text entity may be referred to as the first text mark, and the mark of the fifth text entity may be referred to as the third text mark.
[0197] In another specific implementation, if there is an obstruction between the entity icons of two text entities (which may be referred to as the first text entity and the sixth text entity), the mobile phone may omit the entity icon of one of the text entities, such as omitting the entity icon of the sixth text entity (which may be referred to as the fourth text mark), to avoid obstruction of the entity icon. Exemplarily, in interface 1102, between address 1 marked by location icon 11021 and website 1 marked by network icon 11022, phone 1 is also included. If a phone icon (as shown by the dotted line in interface 1102, which is not actually displayed) is used to mark phone number 1, obstruction between the marks will result, and therefore, the mobile phone may not mark phone 1. This avoids obstruction between marks.
[0198] Specifically, the mobile phone may mark text entities in a top-to-bottom and left-to-right order. If there is an obstruction between the current entity icon and the marked entity icon, the mobile phone may cancel the mark of the current entity icon.
[0199] In another specific implementation, the mobile phone may only mark the number 1 of text entity categories of which the user's interest level is from high to low.
[0200] Specifically, a mobile phone or other device can analyze the interest of a large number of users (or only local users) in multiple types of text entities. Among them, the user's interest in the text entity is positively correlated with the number of times the user triggers the text entity. When a user uses a mobile phone, the more times the user performs a trigger operation (such as clicking, long pressing, etc.) on a certain type of text entity, the higher the user's interest in the text entity.
[0201] In another specific implementation, when using the intelligent recognition function to recognize text, the mobile phone can also determine and mark text entities that match the object entities based on the object entities included in the current interface. Different categories of object entities will result in different categories of marked text entities.
[0202] Specifically, if the first interface includes multiple text entities (including the first text entity and the second text entity), and the current interface includes item entity 1 (recorded as the first item entity), the first text entity that matches item entity 1 is marked from the multiple text entities, and the unmatched second text entity is not marked; if the second interface includes multiple text entities (including the third text entity and the fourth text entity), and the current interface includes item entity 2 (recorded as the second item entity), the third text entity that matches item entity 1 is marked from the multiple text entities, and the unmatched fourth text entity is not marked, and the mark of the third text entity can be called the second text mark.
[0203] That is, the electronic device can mark the text entities of the first category (the entity category of the first text entity and the fourth text entity) in an interface with a certain category (the entity category of the first item entity), but not mark the text entities of the second category (the entity category of the second text entity and the third text entity); and the electronic device can mark the text entities of the second category in an interface with another category (the entity category of the second item entity), but not mark the text entities of the first category. It can be seen that the electronic device can mark matching text entities based on the item entities included in the user interface, thereby providing users with marks for quickly obtaining information from the user interface, which is conducive to improving the efficiency of human-computer interaction.
[0204] Taking the recommended translation function in the third method as an example, after recommending the intelligent recognition function, the mobile phone can display Fig.13A The interface 1301 shown (which can be regarded as a specific first interface) includes a picture 13011 and an entrance 13012 for the smart recognition function. In response to a click operation on the entrance 13012 for the smart recognition function, the mobile phone uses the smart recognition function to recognize the picture 13011, and can recognize the product 13013 (which can be regarded as a specific first item entity). Based on this, the mobile phone can predict that the user may want to view the merchant address and purchase goods, and can determine that the text entities matching the picture 13011 are the address and the website link (which can be regarded as two specific first text entities). Therefore, after the mobile phone uses the smart recognition function to recognize the picture 13011, it can display Fig.13A Interface 1302 is shown. Interface 1302 is a recognition result page. In interface 1302, the mobile phone marks the item entity, such as marking the product 13013 with the item mark 13021; and in interface 1302, the mobile phone also marks the address with a location icon 13022 (which can be regarded as a specific first text mark) so that the user can view the merchant address, and marks the QR code with a scan code icon 13023 (which can be regarded as another specific first text mark) so that the user can scan the code to purchase the product.
[0205] It should be noted that the essence of graphic codes, such as QR codes, is to achieve page jumps, and their function is the same as that of website links. Therefore, mobile phones can also regard graphic codes as a special text entity, and mobile phones can use text extraction functions or intelligent recognition functions to identify graphic codes.
[0206] Furthermore, the mobile phone can mark the text entity that matches the current focus item entity. And, as the focus item entity switches, the marked text entity will also change accordingly. For example, after switching from the first item entity to the third item entity, the mark is also switched from the first text entity to the second text entity, wherein the mark of the second text entity can be called the third text mark. Exemplarily, the current focus item entity is item entity 3, and the mobile phone can mark the text entity that matches item entity 3 from multiple text entities; after the focus item entity switches to item entity 4, the mobile phone can switch to marking the text entity that matches item entity 4. In this way, as the focus item entity switches, the mobile phone can dynamically adjust the marked text entity so that the marked text entity always matches the focus item entity. Among them, regarding the switching of the focus item entity, please refer to the previous introduction on "First, intelligent recognition function", which will not be repeated here.
[0207] In the present application, there is no specific limitation on the content of the mark of each type of text entity (ie, the first text mark or the second text mark).
[0208] In some embodiments, the mobile phone can mark the text entity with the category icon of the text entity. For example, a phone number is marked with a phone icon, a website is marked with a network icon, and an address is marked with a location icon. In this way, the category of the text entity can be clearly indicated by marking.
[0209] In other embodiments, the mobile phone can mark the text entity with the service icon of the service with the highest user interest (which can be called the first service) among the services associated with the text entity. Each type of text entity can be associated with one or more services. A phone number can be associated with multiple services such as making a call, adding a contact, and copying a number. A website can be associated with multiple services such as visiting the website, adding the website to favorites, and sharing the website. An address can be associated with multiple services such as opening in a map, navigating, adding the address to favorites, and sharing the address. A graphic code can be associated with multiple services such as identifying a graphic code, adding the graphic code to favorites, and sharing the graphic code.
[0210] Specifically, a mobile phone or other device can analyze the user's interest in a variety of services. The user's interest in a service is positively correlated with the number of times the service is used. The more times a user selects a service during the use of the mobile phone, the higher the user's interest in the service.
[0211] Take the address-related services as an example, including opening in map, navigation, address collection, and address sharing, which are 4 services in total:
[0212] If the user is most interested in navigation services, the phone will Fig.13A After the interface 1301 shown is identified, it can be displayed Fig. 13BInterface 1311 is shown. In interface 1311, the address is marked using a service icon 13111 of a navigation service.
[0213] If the user is most interested in opening the service in the map, the phone Fig.13A After the interface 1301 shown is identified, it can be displayed Fig. 13B Interface 1312 is shown. In interface 1312, the address is marked using a service icon 13121 that opens the service in the map.
[0214] If the user is most interested in the favorite address, the phone will Fig.13A After the interface 1301 shown is identified, it can be displayed Fig. 13B Interface 1313 is shown. In interface 1313, the address is marked using a service icon 13131 of the favorite service.
[0215] If the user is most interested in the shared address, the phone will Fig.13A After the interface 1301 shown is identified, it can be displayed Fig. 13B Interface 1314 is shown. In interface 1314, the address is marked using a service icon 13141 of the sharing service.
[0216] In the recognition result, the mobile phone will also highlight the text entity. Among them, the mobile phone can highlight the text entity in the form of highlighting, underlining, projection, etc. Exemplarily, in interface 1102 and interface 1302, the address, telephone number, and website are all underlined and projected.
[0217] After obtaining the recognition result, in response to the user's trigger operation (also referred to as the third trigger operation) on the text entity or the mark of the text entity, such as a click operation, the mobile phone can expand the service options of the service associated with the text entity. For ease of explanation, the interface displaying the service options of the text entity can be referred to as the third interface.
[0218] To identify the result page Fig.14 Taking the interface 1401 (which is the same as the interface 1102 in the previous text and will not be described here) as an example, in response to the user clicking on the address 1 in the interface 1401, the mobile phone can display Fig.14 Interface 1402 is shown. Interface 1402 includes pop-up window 14021 and pop-up window 14022. Pop-up window 14021 includes a service option "Open in Map" 140212 for opening a service in a map, a service option "Navigate to" 140213 for a navigation service, a service option "Add to Notes" 140214 for a favorite address service, and a service option "Share" 140215 for a share address service. In addition, pop-up window 14022 includes a route to address 1.
[0219] In this application, there is no specific limitation on the display order of service options.
[0220] In some embodiments, the display order of the service options is always fixed. For example, among the service options of multiple services associated with an address, the display order from front to back is always: the service option of the service to open in the map, the service option of the navigation service, the service option of the address collection service, and the service option of the address sharing service. Fig.14 As shown in interface 1402.
[0221] In other embodiments, the display order of service options matches the user's interest in the services, so that service options for services that the user is most interested in can be displayed first, making it easier for the user to quickly view and operate them.
[0222] In a specific implementation, the mobile phone may display service options in order from high to low according to the user's interest in the service.
[0223] In another specific implementation, the mobile phone may display the service option of the service that the user is most interested in first, and display the service options of other services in a fixed order.
[0224] Combining this implementation with the above-mentioned embodiment of "marking the text entity with the service icon of the service with the highest user interest", after marking the text entity with the service icon of the service with the highest user interest, in response to the user's triggering operation on the text entity, the mobile phone can display the service option of the service indicated by the service icon in the first place, that is, the service option displayed in the first place matches the service icon marked with the text entity in the recognition result. Thus, it is possible to display the service option of the service with the highest user interest in the first place.
[0225] The above description of displaying service options is mainly for the marked address 14011. It should be noted that, in practice, all text entities can be associated with one or more services, not just marked text entities.
[0226] For example, although address 2 in the above interface 1401 is not marked, it can also be associated with multiple services such as making a call, adding a contact, copying a number, etc. In response to the user's click operation on address 2, the mobile phone can also provide service options such as opening in a map, navigating, adding the address to favorites, and sharing the address.
[0227] For unmarked text entities, the display order of service options may also be fixed, or may be matched with the user's interest in the services.
[0228] In the recognition result, in response to a trigger operation on the non-entity text, such as a long press operation, the mobile phone can select the text and provide a shortcut entry to perform a shortcut operation on the non-entity text. For the shortcut entry, please refer to the previous description and will not be repeated here.
[0229] Exemplarily, in response to a user's Fig.14 By long pressing the text "Merchant List" 14011 in the interface 1401, the mobile phone can display Fig.14 Interface 1403 is shown. Different from interface 1401, the text "Merchant List" 14011 is selected in interface 1403, and interface 1403 includes a "Copy" entry 14031, a "Select All" entry 14032, a "Translate" entry 14033, a "Share" entry 14034, and a "Search" entry 14035.
[0230] At this point, it should be noted that: in the process of performing operations on text entities and item entities, the mobile phone can hide the entity mark in the interface. In this way, the interference of the mark can be avoided. For example, in the above interface 904, interface 1001-interface 1004, interface 1402-interface 1403, the entity mark is hidden.
[0231] Third, the translation function.
[0232] The phone uses a translation function that can translate the text in the current interface.
[0233] Taking the method 3 or 4 above as an example, after recommending the translation function, the mobile phone can display Fig.15 The interface 1501 shown in FIG. 1501 (which is the same as the interface 602 described above and will not be described here) includes “Translate” 15011. “Translate” 15011 is the entry for the translation function. In response to the user clicking on “Translate” 15011, the mobile phone may display Fig.15 Interface 1502 is shown. Interface 1502 is a translation result page. Different from interface 1501, the texts in interface 1502 are all in Chinese, that is, all foreign languages are translated into Chinese.
[0234] Furthermore, the translation result page also includes a language selection option for selecting the language of the original text and the translation. Fig.15 Taking the translation result page shown in the middle interface 1502 as an example, the interface 1502 includes a language selection item 15021. In the language selection item 15021, the left side is the original text, and the right side is the translation.
[0235] After obtaining the translation result, the mobile phone can also display the scrolling translation control in the result page. In response to the user's triggering operation on the scrolling translation control, such as a click operation, the mobile phone can scroll the interface content and take a screenshot. That is, a scrolling screenshot. Subsequently, in response to the event of ending the scrolling screenshot, the mobile phone can display the translation result of the scrolling screenshot. Among them, the event of ending the scrolling screenshot includes the event that the scrolling duration reaches duration 2, the event of scrolling to the bottom, or the event of receiving the user's click operation. The user's click operation is used as an example for explanation below. In this way, after the mobile phone translates the interface content currently displayed on the current interface, it can continue to conveniently translate the scrolling interface content.
[0236] by Fig.15 Taking the translation result page shown in the middle interface 1502 as an example, the interface 1502 also includes a "scrolling translation" 15022. The "scrolling translation" 15022 is a scrolling translation control. In response to the user's click operation on the "scrolling translation" 15022 in the interface, the mobile phone can display Fig.15 The interface 1503 shown in FIG. 1503 is scrolling the screenshot. For example, the interface content is scrolling from bottom to top in the direction indicated by the arrow in the interface 1503. In response to the user's click operation in the interface 1503, the mobile phone can display Fig.15 Interface 1504 is shown, and interface 1504 is a translation result page of the screenshot results of the scrolling screenshot.
[0237] After obtaining the translation result of the content of the scrolling screenshot, the mobile phone can also continue to provide a scrolling translation control so as to continue scrolling the screenshot and translating. For example, the interface 1504 includes "continue scrolling" 15041, which is a scrolling translation control.
[0238] Further, after obtaining the translation result of the scrolling screenshot, in response to the user's return operation, the mobile phone can return to the interface of the recommended translation function. Fig.15 By clicking the return control 15042 in the interface 1504, the mobile phone can return to Fig.15 Interface 1501 is shown. In this way, the mobile phone can quickly return to the initial interface before translation, so that the user can operate on the initial interface.
[0239] In some embodiments, the mobile phone can identify whether the current interface is an interface that can be slid up and down. If it is an interface that can be slid up and down, the mobile phone provides a scrolling translation control in the translation result page; if it is not an interface that can be slid up and down, the mobile phone does not provide a scrolling translation control in the translation result page. Among them, the desktop, lock screen interface, and picture viewing interface are usually not interfaces that can be slid up and down. In this way, the mobile phone can dynamically provide scrolling translation controls for different interfaces after obtaining the translation results, thereby ensuring the effectiveness of the displayed controls.
[0240] After obtaining the translation result, in response to the original translation switching operation, such as a click operation on the current interface, the mobile phone can switch to the original text, thereby realizing a quick switch from the translation result to the original text. Fig.15 By clicking on the interface 1502, the mobile phone can switch the Chinese (except the language selection item) in the interface 1502 to English. Fig.15 By clicking on the interface 1504 shown, the mobile phone switches the Chinese language (except the language selection item) in the interface 1504 to English.
[0241] The above introduction to the translation function mainly describes the implementation of the mobile phone's unified translation of the text in the current interface. In practice, after providing an entry for the translation function, the mobile phone can also translate part of the text in the current interface.
[0242] Specifically, after providing the translation function, in response to a user's long press operation on the original text or the translated text, the mobile phone can select the text and provide a translation entry for the selected text. In response to a trigger operation on the translation entry for the selected text, such as a click operation, the mobile phone can display the translation result of the selected text. In this way, after providing the entry for the translation function, the mobile phone can not only translate the entire current interface, but also translate the selected text.
[0243] by Fig.16 For example, the translation result page shown in the middle interface 1601 is used in response to the user's Fig.16 By long pressing the text "have to" in the interface 1601, the mobile phone can display Fig.16 Interface 1602 is shown. The difference from interface 1601 is that the text "have to" in interface 1602 is selected (i.e., "have to" is the selected text), and interface 1602 includes "Translate" 16021. "Translate" 16021 is the translation entry of "have to". In response to the user clicking on "Translate" 16021, the mobile phone can display Fig.16 Interface 1603 is shown. Interface 1603 includes a pop-up window 16031. Pop-up window 16031 includes the translation result of "have to".
[0244] In addition, the mobile phone can provide a translation entry for the selected text as well as quick entry for other text operations, such as Fig.16 The interface 1601 shown includes “Copy” 16022 , “Select All” 16023 , “Share” 16024 , and “Search” 16025 to perform other quick processing on the selected text.
[0245] Fourth, privacy protection function.
[0246] The phone uses a privacy protection function that can block private information in the current interface.
[0247] Taking the above-mentioned method 3 or method 4 as an example, after recommending the privacy protection function, the mobile phone can display Fig.17 The interface 1701 shown in FIG. 1701 (which is the same as the interface 702 described above and will not be described here) includes a “privacy code” 17011. The “privacy code” 17011 is the entrance to the privacy protection function. In response to the user clicking on the “privacy code” 17011, the mobile phone may display Fig.17 Interface 1702 is shown. The difference from interface 1701 is that the email address, ID number, ID picture, address and head portrait in interface 1702 are all coded. That is, the private information is blocked.
[0248] After the private information is blocked, in response to the user clicking on any of the blocked private information, the mobile phone can cancel the blocking of the private information. In response to the user clicking on the private information again, the mobile phone can block the private information again. In this way, the mobile phone can flexibly block or cancel the blocking of private information.
[0249] by Fig.17 Taking the interface 1702 shown as an example, the ID card number 17021 in the interface 1702 is blocked. In response to the user clicking the ID card number 17021 in the interface 1702, the mobile phone can display Fig.17 The interface 1703 shown in FIG. 1702 is different from the interface 1703 in that the ID card number 17021 in the interface 1703 is not blocked. Subsequently, in response to the user clicking the ID card number 17021 in the interface 1703, the mobile phone can display the ID card number 17021 again. Fig.17 The interface 1702 shown is to restore the obscured ID number 17021.
[0250] After the privacy masking is completed, in response to the user's operation to save the masking effect, the mobile phone can also save the masked picture to the gallery. Fig.17 By clicking “Save” 17031 in the interface 1703 shown, the mobile phone can save the picture 17032 displayed in the interface 1703 .
[0251] Furthermore, after the saving is completed, the mobile phone can provide a viewing entrance for the picture. In response to the user's triggering operation on the viewing entrance, such as a click operation, the mobile phone can display the saved picture in the gallery. Fig.17The interface 1704 shown in FIG. 1 includes a prompt “The picture has been saved to the gallery” 17041 and a prompt “View” 17042. “View” 17042 is a viewing entry for the picture. In response to the user clicking on “View” 17042, the mobile phone may display Fig.17 The interface 1705 shown is a viewing interface for pictures in the gallery application. The interface 1705 includes a saved picture 17051, and the privacy information in the picture 17051 is blocked.
[0252] That is to say, with the privacy protection function, the user only needs to trigger the entrance of the privacy protection function and save it once to cover the private information in the current interface and save it as a screenshot, thereby improving the efficiency of human-computer interaction.
[0253] Fifth, select the function.
[0254] The mobile phone adopts the selection function, which can easily process the text in the current interface.
[0255] Taking the above-mentioned method 3 or method 4 as an example, after recommending the selection function, the mobile phone can display Fig.18 The interface 1801 shown in FIG. 1801 (which is the same as the interface 802 described above and will not be described here) includes a "select all" button 18011. The "select all" button 18011 is a control for selecting a function. In response to the user clicking the "select all" button 18011, the mobile phone can select all text in the current interface, such as displaying Fig.18 The interface 1802 shown is different from the interface 1802 in that all the texts in the interface 1802 are in the selected state.
[0256] Furthermore, after selecting all the text, in response to the user adjusting the selection box, the mobile phone can adjust the range of the selected text. Thus, the text can be selected flexibly. Also, after selecting all the text or adjusting the range of the selected text, the mobile phone can also provide shortcuts for various text operations, such as Fig.18 The interface 1802 shown includes “Copy” 18021 , “Favorite” 18022 , “Translate” 18023 , “Share” 18024 and “Search” 18025 to perform quick processing on the selected text.
[0257] The present application also provides an electronic device, which may include: a display screen, a memory, and one or more processors (such as a CPU, a GPU, an NPU, etc.). The display screen, the memory, and the processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device may perform each function or step performed by the device in the above method embodiment.
[0258] The embodiment of the present application also provides a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected by lines. For example, the interface circuit can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit can be used to send signals to other devices (such as a processor). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which are not specifically limited in the embodiment of the present application.
[0259] This embodiment further provides a computer storage medium, in which computer instructions are stored. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the image processing method in the above-mentioned embodiment.
[0260] This embodiment further provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the above-mentioned related steps to implement the image processing method in the above-mentioned embodiment.
[0261] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory so that the chip executes the image processing method in the above-mentioned method embodiments.
[0262] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above and will not be repeated here.
[0263] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0264] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0265] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may be one physical unit or multiple physical units, that is, it may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiment.
[0266] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0267] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program code.
[0268] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, a person of ordinary skill in the art should understand that the technical solution of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present application.
Claims
1. A screen recognition method, characterized in that: Applied to electronic equipment, the method comprises: Displaying a first interface, wherein the first interface includes a first item entity, a first text entity, and a second text entity; In response to a first trigger operation on the first interface, adding a first text mark to the first text entity, and not adding a mark to the second text entity; Displaying a second interface, wherein the second interface includes a second item entity, a third text entity, and a fourth text entity; In response to a second trigger operation on the second interface, adding a second text mark to the third text entity, and not adding a mark to the fourth text entity; Among them, the entity categories of the first item entity and the second item entity are different, the entity categories of the first text entity and the fourth text entity are the same, the entity categories of the second text entity and the third text entity are the same, and the entity categories of the first text entity and the second text entity are different.
2. The method according to claim 1, characterized in that The entity categories of text entities include at least two of the following: address, telephone number, flight information, express delivery number, email address, website link, certificate number for identity identification, and graphic code; The entity categories of an item entity include at least two of the following: animals, plants, buildings, and food.
3. The method according to claim 1 or 2, characterized in that: The first text tag indicates an entity category of the first text entity; or, The first text entity is associated with multiple services, including a first service, the first text tag indicates the first service, the first service is the service most selected by users under a first category of text entities, and the first category is an entity category of the first text entity.
4. The method according to claim 3, characterized in that After adding a first text mark to the first text entity in response to the first trigger operation on the first interface, the method further includes: In response to a third trigger operation on the first text entity or the first text mark, a third interface is displayed, wherein the third interface includes multiple service options, and the multiple service options correspond one-to-one to the multiple services; among the multiple service options, the service option for the first service is displayed first.
5. The method according to any one of claims 1 to 4, characterized in that The first interface also includes a fifth text entity; The method further comprises: In response to the first trigger operation on the first interface, and the entity category of the fifth text entity is different from that of the first text entity, adding a third text tag to the fifth text entity; If the fifth text entity has the same entity category as the first text entity, no mark is added to the fifth text entity.
6. The method according to any one of claims 1 to 5, characterized in that The first interface also includes a sixth text entity; The method further comprises: In response to the first trigger operation on the first interface, and the fourth entity mark of the sixth text entity and the first text mark are not blocked, adding the fourth text mark to the sixth text entity; If the fourth text mark and the first text mark are blocked, no mark is added to the sixth text.
7. The method according to any one of claims 1 to 6, characterized in that The first interface also includes a third item entity; The method further comprises: In response to the first trigger operation on the first interface, highlighting the first item entity and displaying a first shortcut entrance around the first item entity; In response to a fourth trigger operation on the third item entity, the third item entity is highlighted and a second shortcut entrance is displayed around the first item entity.
8. The method according to claim 7, characterized in that In response to the first trigger operation on the first interface, highlighting the first item entity and displaying a first shortcut entrance around the first item entity includes: In response to the first trigger operation on the first interface, and the first item entity meets a first condition, highlighting the first item entity and displaying a first shortcut entrance around the first item entity; The first condition includes at least one of the following: the area of the first item entity is larger than the area of the third item entity, the area blocked by the first item entity is smaller than the area blocked by the third item entity, and the clarity of the edge line of the first item entity is higher than the clarity of the edge line of the third item entity.
9. The method according to claim 7 or 8, characterized in that: The method further comprises: In response to the first trigger operation on the first interface, displaying a first item mark on the third item entity; The fourth trigger operation includes a trigger operation on the first item mark.
10. The method according to any one of claims 7 to 9, characterized in that: The method further comprises: In response to the fourth triggering operation on the third item entity, a third text tag is added to the second text entity.
11. The method according to any one of claims 7 to 10, characterized in that: The method further comprises: In response to a move operation on the highlighted item entity, a fourth interface is displayed, wherein the fourth interface includes a plurality of associated entries, each associated entry corresponds to an application or a service, and the plurality of associated entries include a first associated entry, and the first associated entry corresponds to a first application or a second service; In response to moving the highlighted item entity to the first associated entry, the electronic device displays a fifth interface, where the fifth interface is an interface of the first application or the second service, and the fifth interface includes associated information of the highlighted item entity.
12. The method according to claim 11, characterized in that The step of displaying a fourth interface in response to a move operation on the highlighted item entity comprises: In response to a move operation on the highlighted item entity, move the position of the highlighted item entity in the first interface; In response to the position of the highlighted item entity moving to the target area in the first interface, the fourth interface is displayed.
13. The method according to claim 11 or 12, characterized in that: The highlighted item entity is the first item entity, and the plurality of associated entries include a second associated entry; The highlighted item entity is the third item entity, and the plurality of associated entries include a third associated entry; The second associated entry is different from the third associated entry.
14. The method according to any one of claims 1 to 13, characterized in that The first interface is a camera's viewfinder interface, and the first trigger operation includes a shooting operation.
15. The method according to any one of claims 1 to 13, characterized in that Before adding a first text mark to the first text entity in response to a first trigger operation on the first interface, the method further includes: When the first interface satisfies the second condition, displaying the identification control in the first interface; Among them, the second condition includes: the first interface is a non-blank interface, the first interface includes text entities and / or object entities, and the first trigger operation includes a trigger operation on the identification control.
16. The method according to claim 15, characterized in that When the first interface satisfies the second condition, displaying an identification control in the first interface includes: When the first interface satisfies the second condition, in response to a fifth trigger operation of the user on the first interface, the identification control is displayed in the first interface.
17. The method according to claim 16, characterized in that The method further comprises: In response to a fifth trigger operation of the user on the first interface: When the number of foreign languages in the first interface exceeds a first number, a translation control is displayed in the first interface; wherein, when the number of foreign languages in the first interface does not exceed the first number, the translation control is not displayed, and the translation control is used to trigger the electronic device to translate the foreign language in the first interface; In the case where the first interface includes privacy information, a privacy protection control is displayed in the first interface; wherein, in the case where the first interface does not include privacy information, the privacy protection control is not displayed, and the privacy protection control is used to trigger the electronic device to shield the privacy information in the first interface; When the first interface includes text but does not include an entity, a selection control is displayed in the first interface; wherein, when the first interface does not include text or includes an entity, the selection control is not displayed, and the selection control is used to trigger the electronic device to select the text in the first interface.
18. An electronic device, characterized in that: include: A display screen, one or more processors, and one or more memories; the one or more processors are coupled to the display screen and the one or more memories; the one or more memories are used to store computer program code, the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the electronic device executes the method described in any one of claims 1-17.
19. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method as claimed in any one of claims 1 to 17.
Citation Information
Cited By
Screen recognition method and electronic device
EP4745736A1
Screen recognition method and electronic device
WO2025092139A1