A service recommendation method, an electronic device, and a computer-readable storage medium

By combining the detection and recognition of text entity types and object types, the accuracy problem of electronic devices in identifying information of interest to users is solved, thereby improving the accuracy of service recommendations and user experience.

CN120148040BActive Publication Date: 2026-04-24HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2023-12-06
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, electronic devices are affected by factors such as light, obstruction, dirt, and blur when identifying information of interest to users, resulting in inaccurate identification. Consequently, the services recommended to users are not relevant to the information they are actually interested in, which affects the user experience.

Method used

By combining the entity type of the text with the type of the object in which the text is located, the target service is determined through object detection, text recognition, and entity recognition, which increases the dimensions when determining the target service and improves the accuracy.

Benefits of technology

It improved the accuracy of service recommendations and enhanced the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148040B_ABST
    Figure CN120148040B_ABST
Patent Text Reader

Abstract

The application provides a service recommendation method, an electronic device and a computer readable storage medium, and relates to the technical field of terminals. The method can determine a target service in combination with an entity type of a text and a type of an object where the text is located, increases dimensions referred to when the target service is determined, can make the accuracy of the target service higher, and improves user experience. The method comprises displaying a to-be-processed picture comprising an image of a target object and a text, displaying a target service in response to a first operation of a user on the to-be-processed picture, the target service being associated with a type of the target object and an entity type of a target text of the to-be-processed picture, the target text being a text of the to-be-processed picture, a text region of which overlaps with a region where the image of the target object is located, and the first text being a text of the text of the to-be-processed picture, an entity type of which is a preset entity type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to a service recommendation method, electronic device, and computer-readable storage medium. Background Technology

[0002] In daily life, information that users are interested in is everywhere. For example, this information can exist in electronic images and paper documents. To improve user experience, electronic devices such as mobile phones can identify information of interest in any form and recommend services that match that information. For example, when the user's interest is identified as a tracking number, services such as tracking and queuing for packages can be recommended.

[0003] However, because information of interest is often affected by factors such as light, obstruction, dirt, and blur, the information of interest recognized by electronic devices is inconsistent with the actual information of interest. This further leads to the electronic devices recommending services to users that are not related to the actual information of interest, thus affecting the user experience. Summary of the Invention

[0004] This application provides a service recommendation method, an electronic device, and a computer-readable storage medium for recommending services to users by combining recognized text and the object in which the text is located, which can improve the accuracy of the recommended services and enhance the user experience.

[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, this application provides a service recommendation method, comprising: displaying an image to be processed, the image to be processed including an image of a target object and text; in response to a user's first operation on the image to be processed, displaying a target service, the target service being associated with the type of the target object and the entity type of the target text of the image to be processed, wherein the target text is text in the first text of the image to be processed whose text region overlaps with the area where the image of the target object is located, and the first text is text in the text of the image to be processed whose entity type is a preset entity type.

[0007] Based on the above scheme, electronic devices can determine the target service by combining the entity type of the text and the type of the object in which the text is located. This increases the dimensions considered when determining the target service, making the target service more accurate and improving the user experience.

[0008] In one implementation provided in the first aspect, before displaying the target service, the method further includes: obtaining type information of the target object and entity information of the target text, wherein the type information of the target object is used to indicate a first probability that the target object is each of a variety of types, and the entity information of the target text is used to indicate a second probability that the target text is each of a variety of entity types; determining service recommendation information based on the type information of the target object, the entity information of the target text, and preset configuration information, wherein the configuration information is used to indicate a third probability that text of each entity type appears in each type of object, and the service recommendation information is used to indicate the matching degree between each type of target object and each type of target text; and determining the target service based on the service recommendation information.

[0009] Understandably, when multiple target objects exist, electronic devices can acquire type information of multiple target objects; when multiple target texts exist, electronic devices can acquire entity information of multiple target texts.

[0010] In one implementation provided in the first aspect, obtaining entity information of the target text includes: determining the positional relationship between the image of the target object and the first text based on the positional information of the first text and the positional information of the image of the target object, wherein the positional relationship indicates whether the area where the first text is located overlaps with or does not overlap with the area where the image of the target object is located; and obtaining entity information of the target text based on the positional relationship and the entity information of the first text, wherein the entity information of the first text indicates a second probability that the first text is of each of a variety of entity types. Specifically, determining the positional relationship between the image of the target object and the first text based on the positional information of the first text and the positional information of the image of the target object can accurately determine which target object the first text is located on, thereby identifying the target text.

[0011] In one implementation provided in the first aspect, the method further includes: performing text recognition on the image to be processed to obtain the text and the position information of the text; performing entity recognition on the text to obtain the first text in the text and the entity information of the first text.

[0012] In one implementation provided in the first aspect, obtaining the type information of the target object includes: performing target detection on the image to be processed to obtain the type information of the target object and the position information of the target object's image.

[0013] In one implementation provided by the first aspect, determining the positional relationship between the image of the target object and the first text based on the positional information of the first text and the positional information of the image of the target object includes: zooming in on the area where the first text is located to obtain a candidate region and the positional information of the candidate region; and determining the positional relationship between the image of the target object and the first text based on the positional information of the candidate region and the positional information of the image of the target object.

[0014] Understandably, considering that the text itself occupies a small proportion of the image, in some cases the actual area occupied by the text may exceed the text area determined by the text's location information. Therefore, by expanding the text area determined by the text's location information to obtain candidate areas, the area occupied by the text can be determined more accurately. In this way, the positional relationship between the target object image and the first text is also more accurate.

[0015] In one implementation provided in the first aspect, obtaining entity information of the target text includes: performing text recognition on the image of the target object to obtain all text within the target object; and performing entity recognition on all text within the target object to obtain entity information of the target text. This method eliminates the need to determine the positional relationship between the image of the target object and the target text, simplifying the processing flow of electronic devices and improving the efficiency of recommendation services.

[0016] In one implementation provided in the first aspect, obtaining the type information of the target object includes: performing target detection on the image to be processed to obtain the type information of the target object.

[0017] In one implementation provided in the first aspect, for target objects and target texts that do not overlap in their respective regions, the matching degree between the target object of type i and the target text of type j is positively correlated with the first probability that the target object is of type i, the second probability that the target text is of type j, and the third probability that the text of type j appears under the object of type i.

[0018] In one implementation provided in the first aspect, the type information of the target object, the entity information of the target text, the configuration information, and the service recommendation information satisfy the following:

[0019] M i-j =P1 i *P2 j *P3 i-j

[0020] Among them, P1 i P2 represents the first probability that the target object is of type i. j P3 represents the second probability that the target text is of the j-th entity type. i-j M represents the third probability that text of entity type j will appear under object of type i. i-j Let R represent the matching degree between the target object of type i and the target text of type j, where i ≤ N, j ≤ R, N is the number of multiple types, and R is the number of multiple entity types.

[0021] In one implementation provided in the first aspect, determining the target service based on service recommendation information includes: if it is determined that the entity type of the target text matches the type of the target object based on the service recommendation information, then the service corresponding to the entity type of the target text is determined to be the target service.

[0022] In one implementation provided in the first aspect, the method further includes: if the maximum value in the matching degree is greater than or equal to a preset threshold, determining that the entity type of the target text is the entity type corresponding to the maximum value, determining that the type of the target object is the type corresponding to the maximum value, and determining that the entity type of the target text matches the type of the target object.

[0023] In one implementation provided in the first aspect, the method further includes: modifying the target text according to the type of the target object to improve the accuracy of the text recognition result.

[0024] In a second aspect, this application provides an electronic device, which includes: a memory and one or more processors; wherein the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in the first aspect and any of its implementations.

[0025] Thirdly, this application provides a computer-readable storage medium including computer instructions; when the computer instructions are executed on an electronic device, they cause the electronic device to perform the method as described in the first aspect and any of its implementations.

[0026] The technical effects of any of the design methods in the second to third aspects can be found in the technical effects of different design methods in the first aspect, and will not be repeated here. Attached Figure Description

[0027] Figure 1 A schematic diagram of a scenario provided for an embodiment of this application;

[0028] Figure 2 This is another scenario illustration provided for an embodiment of this application;

[0029] Figure 3 This is another scenario illustration provided by an embodiment of the present application;

[0030] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0031] Figure 5 A layered architecture diagram of an electronic device provided in an embodiment of this application;

[0032] Figure 6This is a schematic diagram illustrating the interaction between modules provided in an embodiment of this application;

[0033] Figure 7 Flowchart of the service recommendation method provided in the embodiments of this application Figure 1 ;

[0034] Figure 8 An example diagram of the first image provided in the embodiments of this application;

[0035] Figure 9 This is a schematic diagram illustrating the positional relationship between the image of the target object and the first text provided in an embodiment of this application.

[0036] Figure 10 A schematic diagram illustrating the positions of the first text and candidate regions provided in this application embodiment;

[0037] Figure 11 Flowchart of the service recommendation method provided in the embodiments of this application Figure 2 ;

[0038] Figure 12 A schematic diagram of a segmented image provided in an embodiment of this application;

[0039] Figure 13 This is a schematic diagram of another segmented image provided in an embodiment of this application. Detailed Implementation

[0040] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0041] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.

[0042] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0043] Currently, mobile phones and other electronic devices can identify information of interest to users in any form and recommend services that match that interest. For example, when a user opens a photo gallery app, the phone can display... Figure 1 The interface 101 shown in (a) is the interface of the gallery application. Interface 101 includes thumbnails of multiple images, such as thumbnail 1011. In response to user interaction with thumbnail 1011, the phone can display... Figure 1 Interface 102 is shown in (b) above. Interface 102 includes an image 103 corresponding to thumbnail 1011. This image 103 may contain information of interest to the user, such as text 1031. For example, if image 103 is a bank card image, the text 1031 of interest to the user could be the bank name, bank card number, etc. In addition, interface 102 also includes a recognition icon 104. In response to the user's operation of the recognition icon 104, the mobile phone can recognize the information of interest to the user, and the mobile phone can display service options matching the recognition result.

[0044] For example, Figure 1(c) shows a service option 105 when the phone misinterprets the text 1031 (actually a bank card number) as a landline number (e.g., 62366813). This service option 105 indicates that the number "62366813" can be called. Specifically, in response to the user clicking service option 105, the phone can call the number "62366813". In response to the user long-pressing service option 105, the phone can display... Figure 1 The service window 106 shown in (d) includes options such as "Call 62366813", "Send Message", and "Add to Contacts". In response to the user's action on the "Call 62366813" option, the phone can call the number "62366813"; in response to the user's action on the "Send Message" option, the phone can display a text message window for the phone and the number "62366813"; in response to the user's action on the "Add to Contacts" option, the phone can add the number "62366813" to the contacts.

[0045] In other possible designs, the phone may also misinterpret the text 1031 as a tracking number, thus displaying service options such as querying or tracking packages.

[0046] It can be seen that the text recognized by the phone is inconsistent with the actual text, and the services recommended by the phone to the user (such as call services) are not related to the actual text. This is because currently, phones only recommend services based on the recognized text. However, text recognized from images is often affected by factors such as lighting, obstruction, dirt, and blurriness, making the recognized text inaccurate. This further leads to electronic devices recommending services to users that are not related to the actual text, affecting the user experience.

[0047] In view of this, this application provides a service recommendation method that can recommend services to users by combining two dimensions: the text itself and the object in which the text is located. This method can improve the matching rate between recommended services and text, and enhance the user experience.

[0048] The method provided in the embodiments of this application can be applied to... Figure 1 The scenario shown, where a user opens the gallery app, can also be applied to screenshotting, photography, and browsing scenarios. See also... Figures 2-3 This is a schematic diagram illustrating an application scenario provided in this application. For example, in the screenshot scenario, such as... Figure 2As shown in (a), the electronic device can display screenshot 201, which may contain information of interest to the user. Taking screenshot 201 as a tracking number detail as an example, the text 2011 that the user is interested in could be the tracking number, address, phone number, etc. In response to the user's two-finger press on screenshot 201, the electronic device can recognize the text 2011 in screenshot 201 and display multiple service options. These multiple service options are matched with the entities recognized by the electronic device from screenshot 201. For example, if screenshot 201 includes three entities: tracking number, address, and phone number, then... Figure 2 As shown in (b), the electronic device can display three service options: “Track Package” 202, “Navigate to XX Community” 203, and “Call 183XXXX1111” 204.

[0049] For example, in a shooting scenario, such as Figure 3 As shown in (a), the electronic device can display a shooting preview interface 301. The captured image 302 displayed in this preview interface 301 may contain information of interest to the user. Taking image 302 as an example of an ID card, the information of interest to the user might include the ID card number, name, address, etc. Furthermore, the shooting preview interface 301 also includes a recognition icon 303. In response to the user's operation of the recognition icon 303, the electronic device can recognize the information of interest in image 302 and display multiple service options. For example... Figure 3 As shown in (b), the mobile phone can display the service option "Navigate to Building X, Apartment X, XX Community, XX City, XX Province" (304).

[0050] The service recommendation method provided in this application can be applied to electronic devices, such as mobile phones, tablets, televisions (also known as smart screens), desktop computers, laptops, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, wearable devices, in-vehicle devices, smart home devices, and / or smart city devices. This application does not impose any special restrictions on the specific type of electronic device.

[0051] Figure 4 A schematic diagram of the hardware structure of an electronic device is shown. (For example...) Figure 4As shown, the electronic device 200 may include: a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0052] The processor 210 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The processor 210 can be the nerve center and command center of an electronic device. The processor 210 can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0053] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0054] In this embodiment of the application, the processor 210 can perform text recognition, entity recognition, object detection, etc. on the image, and determine the service to be recommended to the user based on the corresponding results.

[0055] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0056] The external storage interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0057] Internal memory 221 can be used to store computer executable program code, which includes instructions. Processor 210 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 221. For example, in this embodiment, processor 210 can execute instructions stored in internal memory 221, which may include a program storage area and a data storage area.

[0058] The program storage area can store the operating system, at least one application required for a function (such as a service recommendation function), etc. The data storage area can store data created during the use of the electronic device (such as configuration information, entity information of the target text, location information of the target text, entity information of the first text, location information of the target object, etc., which will be described in detail later). In addition, the internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0059] The charging management module 240 receives charging input from a charger, which can be a wireless charger or a wired charger. While charging the battery 242, the charging management module 240 can also supply power to the electronic device via the power management module 241.

[0060] The power management module 241 connects the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240, and supplies power to the processor 210, internal memory 221, external memory, display 294, camera 293, and wireless communication module 260, etc. In some embodiments, the power management module 241 and the charging management module 240 may also be housed in the same device.

[0061] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor, and baseband processor. In some embodiments, antenna 1 of the electronic device is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, enabling the electronic device to communicate with networks and other devices through wireless communication technology.

[0062] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0063] The mobile communication module 250 can provide solutions for wireless communication applications in electronic devices, including 2G / 3G / 4G / 5G. The mobile communication module 250 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation.

[0064] The mobile communication module 250 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 250 can be housed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 can be housed in the same device.

[0065] The wireless communication module 260 can provide solutions for wireless communication applications in electronic devices, including WLAN (such as wireless fidelity, Wi-Fi) networks, Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, and other wireless communication technologies.

[0066] The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signal, and sends the processed signal to processor 210. The wireless communication module 260 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0067] Electronic devices can implement audio functions through audio modules 270, speakers 270A, receivers 270B, microphones 270C, headphone jacks 270D, and application processors. Examples include music playback and recording.

[0068] Sensor module 280 may include sensors such as pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors. Electronic devices can acquire various data through sensor module 280.

[0069] Electronic devices implement display functions through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0070] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. In this embodiment, display screen 294 can be used to display services determined by processor 210 to be recommended to the user, as well as related interfaces, such as those described above. Figures 1-3 The interface in the middle.

[0071] The electronic device can implement shooting functions through an ISP, camera 293, video codec, GPU, display 294, and application processor. The ISP is used to process data fed back by the camera 293. The camera 293 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras 293, where N is a positive integer greater than 1.

[0072] Buttons 290 include a power button, volume buttons, etc. Buttons 290 can be mechanical buttons or touch-sensitive buttons. Motor 291 can generate vibration alerts. Motor 291 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 292 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 295 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to achieve contact and separation with the electronic device. The electronic device can support one or N SIM card interfaces, where N is a positive integer greater than 1. SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.

[0073] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments, the electronic device may include more or fewer modules than those provided in the above embodiments, and the modules may employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0074] The software system of the aforementioned electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture of the Android system as an example to exemplify the software structure of the electronic device.

[0075] like Figure 5 As shown, a layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through interfaces. In some embodiments, the Android system may include an application layer, an application framework layer, an Android runtime and algorithm library, a hardware abstraction layer (HAL), and a kernel layer. It should be noted that this application uses the Android system as an example; however, the solution can also be implemented in other operating systems (such as iOS), provided that the functions implemented by each module are similar to those in the embodiments of this application.

[0076] The application layer can include a series of applications. For example, the application layer may include applications such as camera, gallery, note-taking, document editing, navigation, and service association, without specific limitations. Service association applications can provide text recognition and recommendation services to other applications (such as camera and gallery applications).

[0077] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions. For example, it may include a window manager, activity manager, resource manager, notification manager, camera framework, etc., but this application embodiment does not impose any limitations on this.

[0078] The algorithm library can include multiple functional modules. For example, there are object detection modules, text recognition modules, entity recognition modules, and positional relationship determination modules. The object detection module detects the presence of objects of a preset type in an image and determines the probability of each of the various object types. The text recognition module identifies text in an image and determines its location. The entity recognition module determines whether the text belongs to a preset entity type and the probability of each entity type. The positional relationship determination module determines the positional relationship between the objects and text, indicating whether the areas containing the text and the objects overlap.

[0079] Alternatively, this algorithm library can also be called the algorithm engine layer.

[0080] The Android runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core libraries consist of two parts: one part contains the functionalities that Java needs to call, and the other part contains the core Android libraries. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0081] The HAL layer is a wrapper around Linux kernel drivers, providing interfaces to the upper layers and shielding them from the implementation details of the underlying hardware.

[0082] The HAL layer can include camera HAL, sensor HAL, audio HAL, etc.

[0083] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, audio drivers, and camera drivers.

[0084] The following describes the software modules and interactions between modules involved in the service recommendation method provided in the embodiments of this application. For example... Figure 6 As shown, the camera application at the application layer can send image detection requests to service-related applications at the same layer. The service-related applications can send object detection requests to the object detection module and text recognition module of the algorithm library, respectively. The object detection module performs object detection on the image and obtains the object detection result, which includes the probability that the object is of multiple types. The text recognition module performs text recognition on the image and obtains the text recognition result, which includes the text and its location information. The text recognition module can also input the text recognition result into the entity recognition module, which identifies entities in the text and the probability that the text is of multiple entity types. The object detection module can send the object detection result to the positional relationship determination module, and the entity recognition module can send the probability that the text is of multiple entity types and the text's location information to the positional relationship determination module. Thus, the positional relationship determination module can determine the positional relationship between the text and the object. Then, the positional relationship determination module sends the positional relationship between the text and the object, the probability that the object is of multiple types, and the probability that the text is of multiple entity types to the service-related applications. The service-related applications combine this information with preset configuration information to determine the matching degree between the text and the object, and then recommend services based on the matching degree.

[0085] For ease of understanding, the service recommendation method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0086] See Figure 7 The following is a flowchart illustrating the service recommendation method provided in the embodiments of this application. Figure 1 .like Figure 7 As shown, the recommended methods for this service include S501 to S509.

[0087] S501, retrieve the first image.

[0088] The first image can be an image pre-stored by the electronic device, for example... Figure 1 The image 103 shown can be an image captured in real time by an electronic device, such as an image captured by an electronic device and displayed in a shooting preview interface. Figure 3 The image 302 shown can be an image sent from other devices, such as an image received by a user through a chat application while chatting with other users using an electronic device; no specific limitation is made here. Optionally, this first image can also be referred to as the image to be processed.

[0089] In one possible design, on the one hand, the first image may include an image of the target object. The target object refers to an object of a preset type, which includes, but is not limited to, express delivery slips, bank cards, ID cards, business cards, books, screenshots, etc. On the other hand, the first image may include text. This text may include Chinese characters, numbers, symbols, etc., without specific limitations.

[0090] For example, in such Figure 8 The bank card 401 and ID card 402 in the first image shown are both target objects. This first image includes an image of bank card 401 (the image area marked by the dotted frame around bank card 401 in the image) and an image of ID card 402 (the image area marked by the dotted frame around ID card 402 in the image). Additionally, the first image also includes text such as "XX Bank", "62236681XXXXXXXXXXX1", "ATM", "Name Chen XX", "Address XX Province XX City XX District XX Community XX Building X Number", and "Citizen ID Card Number 5XXXXXXXXXXXXXXXXXX".

[0091] For example, Figure 1 In the image 103 shown, the bank card is the target object, and the image of the area where the bank card is located is the image of the target object; Figure 2 The screenshot 201 shown is the target object, and therefore screenshot 201 itself is the image of the target object; Figure 3 The ID card in the image 302 shown is the target object, and the image of the area where the ID card is located is the image of the target object.

[0092] S502, Perform target detection on the first image to obtain the type information of the target object in the first image and the position information of the target object in the image.

[0093] The target object type information is used to indicate the probability (also known as the first probability) that the target object is of each of several types. In the embodiments of this application, the various types include express delivery slips, bank cards, ID cards, business cards, books, screenshots, etc.

[0094] In one possible design, the electronic device can input a first image into an object detection model. This object detection model is a pre-trained, convergent model capable of detecting target objects in the image. The input to the object detection model is the image, and the output of the object detection model is the location information and type information of each target object in the image. Optionally, Figure 5 The target detection module in the system can integrate the above-mentioned target detection model.

[0095] For example, for Figure 8The first image shows that the electronic device can obtain the type information of bank card 401 and ID card 402. The type information of bank card 401 includes the probability that it represents a delivery slip, bank card, ID card, business card, or book. The type information of ID card 402 includes the probability that it represents a delivery slip, bank card, ID card, business card, or book.

[0096] The location information of the target object's image is used to indicate the position of the target object's image in the first image.

[0097] In one possible design, the positional information of the target object's image can include the coordinates of the four vertices of the target object's image. For example, Figure 8 As shown, the coordinate system in question uses the top-left vertex of the first image as its origin, the width of the first image as its x-axis, and the height of the first image as its y-axis. It should be noted that, unless otherwise specified, all coordinates appearing in the following text will be referenced to this coordinate system.

[0098] In one possible design, if the electronic device determines that there is no target object in the first image after performing target detection, the electronic device can end the process.

[0099] S503, perform text recognition on the first image to obtain the text in the first image and the text's location information.

[0100] In this embodiment, the electronic device can use an optical character recognition (OCR) algorithm to perform text recognition on the first image, obtaining the text in the first image and the text's location information. Optionally, Figure 5 The text recognition module in the system can integrate the aforementioned OCR algorithm.

[0101] It should be noted that the text can be any text in the first image, or the text selected by the user in the first image; there are no specific restrictions here.

[0102] For example, electronic devices to Figure 8 Text recognition of the first image shown can yield text such as "XX Bank", "62236681XXXXXXXXXXX1", "ATM", "Name Chen XX", "Address XX Province XX City XX District XX Community XX Building X Number", "Citizen ID Number 5XXXXXXXXXXXXXXXXX", as well as the location information of each text.

[0103] The text's location information is used to indicate the text's position within the first image. In one possible design, the text's location information could include the coordinates of the four vertices of the area where the text is located.

[0104] In this embodiment, the text in the first image includes first text. The first text is text in the first image whose entity type is a preset entity type. The preset entity type is any one of the following: express delivery tracking number, mobile phone number, landline number, service number, address, bank card number, ID card number, email address, website address, and name.

[0105] For example, the texts such as “62236681XXXXXXXXXXX1”, “Name Chen XX”, “Address XX Province XX City XX District XX Community XX Building X Number”, and “Citizen ID Number 5XXXXXXXXXXXXXXXXX” are all first texts, while texts such as “XX Bank” and “ATM” are not first texts.

[0106] It should also be noted that there is no strict order of execution between S502 and S503. S502 can be executed first and then S503, or S503 can be executed first and then S502, or both S502 and S503 can be executed simultaneously. No specific restrictions are imposed here.

[0107] S504, perform entity recognition on the text in the first image to obtain the entity information of the first text.

[0108] The entity information of the first text is used to indicate the probability (also known as the second probability) that the first text in the first image belongs to each of a variety of entity types. These various entity types include tracking number, mobile phone number, landline number, service number, address, bank card number, ID card number, email address, website address, and name.

[0109] For example, Figure 8 In the first image shown, the entity type of "62236681XXXXXXXXXXX1" is a bank card number, the entity type of "Name Chen XX" is a name, the entity type of "Address XX Province XX City XX District XX Community XX Building X Number" is an address, and the entity type of "Citizen ID Number 5XXXXXXXXXXXXXXXXX" is an ID number. The entity types of texts such as "ATM" and "XX Bank" do not belong to any of the above entity types. Therefore, the electronic device can determine that the texts "62236681XXXXXXXXXXX1", "Name Chen XX", "Address XX Province XX City XX District XX Community XX Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX" are all first texts, while the texts "ATM" and "XX Bank" are not first texts.

[0110] In one possible design, the electronic device can input the text from the first image obtained in S503 into a named entity recognition (NER) model. This NER model is a pre-trained model used to extract entities from text. The input to the NER model is text, and the output of the NER model includes whether each piece of text is the first text, and the probability that the first text belongs to each of several entity types. In the embodiments of this application, Figure 5 The entity recognition module shown can integrate the aforementioned NER model.

[0111] For example, for Figure 8 The first image shows that an electronic device can input text such as "XX Bank", "62236681XXXXXXXXXXX1", "Name Chen XX", "Address XX Province XX City XX District XX Community Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX" into a NER model. The model then calculates the probabilities that "62236681XXXXXXXXXXX1", "Name Chen XX", "Address XX Province XX City XX County XX Village", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX" represent, respectively, a tracking number, mobile phone number, landline number, service number, address, bank card number, ID number, email address, address, website, and name. The NER model can determine that the entity type of "XX Bank" does not belong to any of the above entity types, and therefore will not output the probability that "XX Bank" represents a tracking number, mobile phone number, landline number, service number, address, bank card number, ID number, email address, address, website, or name.

[0112] In one possible design, if the electronic device determines that the first text does not exist in the first image after performing entity recognition on the text of the first image, the electronic device can end the process.

[0113] S505, Filter the location information of the first text from the location information of the text to obtain the location information of the first text.

[0114] Understandably, given that the location information of each text has been determined, and that the texts that are the first text have been identified, the electronic device can directly filter out the location information of the first text from the text location information.

[0115] S506, The positional relationship between the image of the target object and the first text is obtained based on the positional information of the image of the target object and the positional information of the first text.

[0116] In this embodiment, the positional relationship between a target object and a first text can include two types: far away and not far away. Wherein, a positional relationship of far away means that the first text is located outside the image of the target object, or it can be understood as the area where the first text is located and the area where the target object's image is located do not overlap. A positional relationship of not far away means that part or all of the first text is located within the area where the target object's image is located, or it can be understood as part or all of the area where the first text is located overlaps with the area where the target object's image is located.

[0117] It should be noted that the positional relationship between the image of the target object and the first text described in the embodiments of this application includes the positional relationship between any target object and any first text. For example, if there is one target object and one first text in the first image, the electronic device can directly obtain the positional relationship between the target object and the first text. As another example, if there are target object 1, target object 2, first text 1, first text 2, and first text 3 in the first image, the positional relationship between the image of the target object and the first text includes the positional relationships between target object 1 and first text 1, first text 2, and first text 3, respectively, and the positional relationships between target object 2 and first text 1, first text 2, and first text 3, respectively. In this way, the electronic device can determine which target object the first text is specifically located on.

[0118] In one possible design, for any given first text and any given target object, the electronic device can determine whether all four vertices of the region containing the first text are located outside the region containing the target object's image. If all four vertices of the region containing the first text are located outside the region containing the target object's image, it indicates that the region containing the first text and the region containing the target object's image do not overlap, thus determining that the positional relationship between the first text and the target object's image is far apart. If at least one of the four vertices of the region containing the first text is located within the region containing the target object's image, it indicates that the region containing the first text and the region containing the target object's image partially overlap, thus determining that the positional relationship between the first text and the target object's image is not far apart.

[0119] See Figure 9This diagram illustrates the positional relationship between the image of the target object and the first text. Specifically, all four vertices of the region containing the first text 602 are located within the region containing the image of the target object 601, indicating that the positional relationship between the first text 602 and the image of the target object 601 is not far apart. Two vertices of the region containing the first text 603 are located within the region containing the image of the target object 601, and two vertices are located outside the region containing the image of the target object 601, indicating that the positional relationship between the first text 603 and the image of the target object 601 is not far apart. All four vertices of the region containing the first text 604 are located outside the region containing the image of the target object 601, indicating that the positional relationship between the first text 604 and the image of the target object 601 is far apart.

[0120] In this embodiment, the electronic device can determine whether the vertices of the first text are located within the image of the target object based on the coordinates of each vertex of the first text and the coordinates of the four vertices of the target object's image. For example, as shown... Figure 9 As shown, for any vertex P in the first text, the coordinates of vertex P are (X, Y). The location information of the target object includes the coordinates of the four vertices of the target object: (X1, Y1), (X2, Y2), (X3, Y3), and (X4, Y4). The electronic device can determine Xmax, Xmin, Ymax, and Ymin among the coordinates of the four vertices. Xmax is the maximum value of the x-coordinate among the four vertices, Xmin is the minimum value of the x-coordinate among the four vertices, Ymax is the maximum value of the y-coordinate among the four vertices, and Ymin is the minimum value of the y-coordinate among the four vertices. If Xmin ≤ X ≤ Xmax and Ymin ≤ Y ≤ Ymax, it means that vertex P is inside the target object; if X < Xmin or X > Xmax or Y < Ymin or Y > Ymax, it means that vertex P is outside the target object.

[0121] For example, such as Figure 9 As shown, the position information of the target object includes the coordinates of the lower left vertex (X1, Y1), the upper left vertex (X2, Y2), the upper right vertex (X3, Y3), and the lower right vertex (X4, Y4). Here, Xmin is X1, Xmax is X3, Ymin is Y2, and Ymax is Y4. If the coordinates of vertex P satisfy X1≤X≤X3 and Y2≤Y≤Y4, then vertex P is determined to be inside the target object; otherwise, vertex P is determined to be outside the target object.

[0122] In one possible design, the electronic device can determine whether the four vertices of the first text are within the target object using the following code.

[0123] public boolean contains(X,Y){

[0124] return Xmin<Xmax&&Ymin<Ymax&&X≥Xmin&&X≤Xmax&&Y≥Ymin&&Y≤Ymax

[0125] }

[0126] In another possible design, the electronic device can first zoom in on the area containing the first text to obtain the location information of the candidate region. Then, for this candidate region, the electronic device can determine whether all four vertices of the candidate region are outside the target object. If all four vertices of the candidate region are outside the target object, the positional relationship between the first text and the target object can be determined to be far apart; if at least one of the four vertices of the candidate region is inside the target object, the positional relationship between the first text and the target object can be determined to be not far apart.

[0127] For example, such as Figure 10 As shown, the electronic device expands the first text area 701 outwards at both ends of its height by 'a' and outwards at both ends of its width by 'b', based on the first text area 701, to obtain the candidate area 702.

[0128] Taking the coordinates of the four vertices of the candidate region as (X5, Y5), (X6, Y6), (X7, Y7), and (X8, Y8), and the coordinates of the four vertices of the first text as (X9, Y9), (X10, Y10), (X11, Y11), and (X12, Y12), the position information of the candidate region and the position information of the first text satisfy the following:

[0129] (1) The squared distance between the two corresponding vertices in the region where the first text is located and the candidate region is a fixed value:

[0130] (X9-X5) 2 +(Y9-Y5) 2 =a 2 +b 2 (X10-X6) 2 +(Y10-Y6) 2 =a 2 +b 2 ,

[0131] (X11-X7) 2 +(Y11-Y7) 2 =a 2 +b 2 (X12-X8) 2 +(Y12-Y8) 2 =a 2 +b 2

[0132] (2) The slopes of the two corresponding edges in the region where the first text is located and the candidate region are equal:

[0133] Where the bounding boxes containing the first text region and the candidate region are not parallel to the coordinate system:

[0134] (Y6-Y5) / (X6-X5)=(Y10-Y9) / (X10-X9);

[0135] (Y7-Y6) / (X7-X6)=(Y11-Y10) / (X11-X10);

[0136] (Y8-Y7) / (X8-X7)=(Y12-Y11) / (X12-X11);

[0137] (Y8-Y5) / (X8-X5)=(Y12-Y9) / (X12-X9);

[0138] When the rectangle containing the first text and the candidate region is parallel to the coordinate system:

[0139] X5=X6, X10=X9=X5+b, X7=X8, X12=X11=X7-b.

[0140] (3) The distances from the four vertices of the candidate region to the corresponding edges of the region where the first text is located are fixed values:

[0141] Taking vertex (X6, Y6) of the candidate region as an example, the region where the first text is located includes edge 1 formed by vertices (X10, Y10) and (X11, Y11). Therefore, the distance from vertex (X6, Y6) to edge 1 satisfies the following formula:

[0142]

[0143] Similarly, the region where the first text is located also includes edge 2 formed by vertices (X11,Y11) and (X8,Y8), edge 3 formed by vertices (X8,Y8) and (X9,Y9), and edge 4 formed by vertices (X9,Y9) and (X10,Y10). Calculate the distance from vertex (X7,Y7) of the candidate region to edge 2 of the region where the first text is located, calculate the distance from vertex (X8,Y8) of the candidate region to edge 3 of the region where the first text is located, and calculate the distance from vertex (X5,Y5) of the candidate region to edge 4 of the region where the first text is located, and list the corresponding formulas.

[0144] Based on the three sets of formulas above, the electronic device can calculate a and b, and then determine the position information of the candidate region based on the position information of the first text and a and b. Then, the electronic device can further determine the positional relationship between the candidate region and the target object image based on the position information of the candidate region and the target object image, to determine whether the positional relationship between the first text and the target object image is far away or not far away.

[0145] Understandably, considering that the text itself occupies a small proportion of the image, in some cases the actual area occupied by the text may exceed the text area determined by the text's location information. Therefore, by expanding the text area determined by the text's location information to obtain candidate areas, the area occupied by the text can be determined more accurately.

[0146] Optionally, the electronic device can further expand the region where the image of the target object is located to obtain a candidate region, and then determine whether all four vertices of the region where the first text is located are outside the candidate region. If all four vertices of the region where the first text is located are outside the candidate region, the positional relationship between the first text and the target object can be determined to be far apart; if at least one of the four vertices of the region where the first text is located is located within the candidate region, the positional relationship between the first text and the target object can be determined to be not far apart. Alternatively, the electronic device can expand both the region where the first text is located and the region where the image of the target object is located, and then determine the positional relationship between the first text and the target object based on the two expanded regions.

[0147] S507, Obtain the entity information of the target text based on the positional relationship between the image of the target object and the first text, as well as the entity information of the first text.

[0148] Among them, the target text is the text in the first text of the first image that overlaps with the area where the image of the target object is located.

[0149] The entity information of the target text is used to indicate the probability (also known as the second probability) of the target text being each of a variety of entity types.

[0150] Understandably, the electronic device has already acquired the probability that the first text belongs to each of the multiple entity types (the entity information of the first text mentioned above). Therefore, the electronic device can determine which specific target texts are included in the image of each target object based on the positional relationship between the image of the target object and the first text, and thus obtain the entity information of the target text.

[0151] For example, continue with Figure 8Taking the first image as an example, as explained earlier, the text in this image, including "62236681XXXXXXXXXXX1", "Name Chen XX", "Address XX Province XX City XX District XX Community Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX", is the first text. Since the area containing "62236681XXXXXXXXXXX1" overlaps with the area containing the image of bank card 401, the electronic device can determine that "62236681XXXXXXXXXXX1" is located within the image of bank card 401. Similarly, since the areas containing "Name Chen XX", "Address XX Province XX City XX District XX Community Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX" overlap with the area containing the image of ID card 402, the electronic device can determine that "Name Chen XX", "Address XX Province XX City XX District XX Community Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX" are located within the image of ID card 402. Furthermore, the electronic device can determine that "62236681XXXXXXXXXXX1", "Name Chen XX", "Address XX Province XX City XX District XX Community Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX" are all target text. Therefore, the entity information of the target text, including "62236681XXXXXXXXXXX1", "Name Chen XX", "Address XX Province XX City XX District XX Community Building X Number", and "Citizen ID Number 5XXXXXXXXXXXXXXXXX", represents the probability of each of the various entity types.

[0152] S508 determines service recommendation information based on the type information of the target object, the entity information of the target text, and the preset configuration information.

[0153] The preset configuration information indicates the probability (also known as the third probability) of text of each entity type appearing under each type of object. It should be noted that each third probability in this configuration information can be obtained through statistics or experiments by the developers, representing the probability of various different entity types appearing under different types in actual applications.

[0154] For example, the preset configuration information can be shown in Table 1.

[0155] Table 1

[0156]

[0157]

[0158] For example, Table 1 indicates that the probability of a tracking number appearing on a waybill is 95%, the probability of a mobile phone number appearing on a waybill is 80%, the probability of a bank card number appearing on a waybill is 5%, and the probability of an ID card number appearing on a waybill is 5%. Table 1 also indicates that the probability of a tracking number appearing on a bank card is 5%, the probability of a mobile phone number appearing on a bank card is 5%, the probability of a bank card number appearing on a bank card is 95%, and the probability of an ID card number appearing on a bank card is 5%.

[0159] It should be noted that the types, entity types, and third probabilities shown in Table 1 above are only illustrative. In other embodiments, the configuration information may include more types and entity types than those in Table 1, and the third probabilities may also be adjusted.

[0160] Service recommendation information is used to indicate the degree of matching between each type of target object and each type of target text.

[0161] In this embodiment, for target objects and target text that are not far apart in location, the matching degree between the i-th type of target object and the j-th type of entity text is positively correlated with the following three parameters: a first probability that the target object is of type i, a second probability that the target text is of type j, and a third probability that text of type j appears under the object of type i. Where i ≤ N, j ≤ R, N is the number of multiple types, and R is the number of multiple entity types.

[0162] In one possible design, the target object's type information, the target text's entity information, configuration information, and service recommendation information satisfy the following:

[0163] M i-j =P1 i *P2 j *P3 i-j

[0164] Among them, P1 i P2 represents the first probability that the target object is of type i. j P3 represents the second probability that the target text is of the j-th entity type. i-j M represents the third probability of the j-th entity type occurring under the i-th type. i-j This represents the degree of matching between the target object of type i and the target text of type j.

[0165] The following text will take the first image, which includes a target object and a target text, as an example to illustrate the process of determining service recommendation information based on the type information of the target object, the entity information of the target text, and the preset configuration information.

[0166] For example, the target object's type information indicates a 70% probability that the target object is a waybill, a 5% probability that it is a bank card, a 7% probability that it is an ID card, a 13% probability that it is a business card, and a 5% probability that it is a book. The target text's entity information indicates a 30% probability that the target text is a waybill number, a 6% probability that it is a mobile phone number, a 6% probability that it is a landline number, an 8% probability that it is a service number, a 3% probability that it is an address, a 40% probability that it is a bank card number, a 3% probability that it is an ID card number, a 1% probability that it is an email address, a 2% probability that it is a website address, and a 1% probability that it is a name.

[0167] According to Table 1, the probability of a tracking number appearing on a waybill is 95%. Therefore, the matching degree when the target object is a waybill and the target text is a tracking number is 70% × 30% × 95% = 19.95%.

[0168] According to Table 1, the probability of a mobile phone number appearing on a waybill is 80%. Therefore, the matching degree when the target object is a waybill and the target text is a mobile phone number is 70% × 6% × 80% = 3.36%.

[0169] According to Table 1, the probability of a landline number appearing on a waybill is 50%. Therefore, the matching degree when the target object is a waybill and the target text is a landline number is 70% × 6% × 50% = 2.10%.

[0170] According to Table 1, the probability of a service number appearing on a waybill is 50%. Therefore, the matching degree when the target object is a waybill and the target text is a service number is 70% × 8% × 50% = 2.80%.

[0171] According to Table 1, the probability of an address appearing on a waybill is 80%. Therefore, the matching degree when the target object is a waybill and the target text is an address is 70% × 3% × 80% = 1.68%.

[0172] According to Table 1, the probability of a bank card number appearing on a waybill is 5%. Therefore, the matching degree when the target object is a waybill and the target text is a bank card number is 70% × 40% × 5% = 1.40%.

[0173] According to Table 1, the probability of an ID number appearing on a waybill is 5%. Therefore, the matching degree when the target object is a waybill and the target text is an ID number is 70% × 3% × 5% = 0.11%.

[0174] According to Table 1, the probability of an email address appearing on a waybill is 50%. Therefore, the matching degree when the target object is a waybill and the target text is an email address is 70% × 1% × 50% = 0.35%.

[0175] According to Table 1, the probability of a URL appearing on a waybill is 30%. Therefore, the matching degree when the target object is a waybill and the target text is a URL is 70% × 2% × 30% = 0.42%.

[0176] According to Table 1, the probability of a name appearing on a delivery slip is 80%. Therefore, the matching degree when the target object is a delivery slip and the target text is a URL is 70% × 1% × 80% = 0.6%.

[0177] Similarly, electronic devices can obtain service recommendation information as shown in Table 2.

[0178] Table 2

[0179]

[0180] When the first image includes multiple target objects and multiple target texts, the electronic device can determine the degree of matching between any two target objects and target texts that are not far apart in position, according to the method described above.

[0181] For example, such as Figure 8 As shown, the first image includes two target objects: a bank card (401) and an ID card (402). Specifically, area 301 includes target text 1 (e.g., 62236681XXXXXXXXXXX1), and area 302 includes target text 2 (e.g., “Chen XX”), target text 3 (e.g., “XX Province XX City XX District XX Community XX Building X Number”), and target text 4 (e.g., “5XXXXXXXXXXXXXXXXX”).

[0182] The electronic device can then obtain the type information of bank card 401, the type information of ID card 402, the entity information of target text 1, the entity information of target text 2, and the entity information of target text 3.

[0183] Then, since the positional relationship between target text 1 and bank card 401 is not far apart, and the positional relationship between target text 2 and target text 3 and ID card 402 is not far apart, the electronic device can determine the matching degree between bank card 401 and target text 1 based on the type information of bank card 401 and the entity information of target text 1, determine the matching degree between ID card 402 and target text 2 based on the type information of ID card 402 and the entity information of target text 2, and determine the matching degree between ID card 402 and target text 3 based on the type information of ID card 402 and the entity information of target text 3, thereby obtaining service recommendation information.

[0184] S509, recommending services based on service recommendation information.

[0185] Among them, electronic devices can determine the target service based on service recommendation information and display the target service, thereby realizing service recommendation.

[0186] In this embodiment, the electronic device can determine the maximum value in the matching degree based on service recommendation information. If the maximum value in the matching degree is greater than or equal to a preset threshold, it can be determined that the type corresponding to the maximum value is the type of the target object, the entity type corresponding to the maximum value is the entity type of the target text, and the target object and the target text match. Therefore, the service corresponding to the entity type of the target text is determined as the target service.

[0187] The preset threshold can be set according to actual needs, and no specific restrictions are imposed here.

[0188] For example, as shown in Table 2, the maximum matching degree is the matching degree between the target object of type "express delivery slip" and the target text of entity type "express delivery tracking number", which is 19.95%. If this matching degree is greater than or equal to the preset threshold, it can be determined that the type of the target object is "express delivery slip" and the entity type of the target text is "express delivery tracking number". Therefore, the electronic device determines that the target service is a service related to express delivery tracking number.

[0189] In one possible design, if the target text is a tracking number, the target service can be a tracking service; if the target text is a mobile phone number, the target service can be a service for making calls, creating contacts, etc.; if the target text is an address, the target service can be a navigation service; if the target text is a bank card number, the target service can be a money transfer service; if the target text is a website address, the target service can be a webpage search service, etc., without specific limitations.

[0190] If the maximum value of the matching degree corresponding to the target text is less than the preset threshold, it indicates that the target object and the target text corresponding to the matching degree do not match. The electronic device may not recommend any service, or recommend general services such as copying and bookmarking, without specific restrictions.

[0191] It should be noted that if the entity types of multiple target texts can be determined based on service recommendation information, the electronic device can recommend services corresponding to each of these target text entity types. Alternatively, the electronic device can recommend services corresponding to only one or a few target text entity types; no specific restrictions are imposed here.

[0192] For example, with Figure 2Taking the scenario shown as an example, in response to the user's two-finger press on screenshot 201, the electronic device can perform object detection and text recognition on screenshot 201. Specifically, the electronic device can determine that screenshot 201 is a target object and obtain the probability that the target object is one of the following types: express delivery slip, bank card, ID card, business card, book, etc., as well as the text in screenshot 201, such as "78744071XXXXXXX", "XX Community", "183XXXX1111", "Shipped", etc., and the location information of all text. Then, the electronic device can perform entity recognition on the text such as "78744071XXXXXXX", "XX Community", "183XXXX1111", "Shipped", etc., obtaining the probability that "78744071XXXXXXX", "XX Community", and "183XXXX1111" are each of the various entity types. Next, the electronic device can filter from the location information of all text to obtain the location information of the first text, and obtain the positional relationship between the first text and the target object based on the location information of the first text and the location information of the target object image. In this scenario, all first texts are not far from the target object, therefore all first texts are target texts. Next, the electronic device can determine the following based on the probability that the target texts such as "78744071XXXXXXX", "XX Community", and "183XXXX1111" belong to each of several entity types, the probability that the target object belongs to each of several types, and preset configuration information: the matching degree between each type of target object and each type of target text "78744071XXXXXXX", the matching degree between each type of target object and each type of target text "XX Community", and the matching degree between each type of target object and each type of target text "183XXXX1111". Furthermore, based on the above matching degrees, it determines that the entity type of the target text "78744071XXXXXXX" is a tracking number, the entity type of the target text "XX Community" is an address, and the entity type of the target text "183XXXX1111" is a mobile phone number. Therefore, electronic devices can recommend services corresponding to tracking numbers, such as package tracking; services corresponding to addresses, such as navigation; and services corresponding to phone numbers, such as calling services.

[0193] Optionally, the recommended service for electronic devices can take the form of displaying relevant service options on the interface. For example, such as... Figure 2 As shown in (b), service recommendations are made through three service options: “Track Package” 202, “Navigate to XX Community”, and “Call 183XXXX1111” 204.

[0194] In other embodiments, the electronic device may recommend only one or two of the following services: package tracking, navigation, or call services; no specific limitations are imposed. Furthermore, the form of service recommendation is not limited to this. Figures 1-3 The form shown can also be represented in other forms.

[0195] As can be seen, this application considers not only the text content itself, but also the type of the object in which the text is located when recommending services. The service corresponding to the text will only be recommended when the matching degree between the text and the object is higher than a threshold. This approach helps to improve the accuracy of the recommended services and enhance the user experience.

[0196] In one possible design, the electronic device can also correct target text that is not far from the target object in terms of its positional relationship to the target object, based on the type of the target object. This approach helps improve the accuracy of text recognition results.

[0197] For example, an electronic device determines that the target text is "51XXXo1996XXXX11o5", whose entity type is an ID card number and the target object type is an ID card. Considering that the probability of the letter "o" appearing in an ID card scenario is extremely low, the electronic device can change the letter "o" to the number "0", thereby correcting the target text to "51XXX01996XXXX1105".

[0198] In the above embodiments, the electronic device performs object detection and text recognition on the first image respectively, which allows the electronic device to execute the object detection and text recognition processes in parallel, thereby improving processing efficiency.

[0199] In another possible design, embodiments of this application provide an alternative service recommendation method that does not require determining the positional relationship between the target text and the image of the target object.

[0200] See Figure 11 The following is a flowchart illustrating the service recommendation method provided in the embodiments of this application. Figure 2 .like Figure 11 As shown, the recommended method for this service includes S801 to S806. It should be noted that the embodiments of this application and the above... Figure 7 The process and principle of the illustrated embodiment are similar, and will not be repeated in this embodiment. Please refer to the previous description.

[0201] S801, retrieve the first image.

[0202] For a description of the first image, please refer to S501, which will not be repeated here.

[0203] S802, Perform target detection on the first image to obtain the type information of the target object in the first image and the image of the target object.

[0204] In one possible design, the object detection model can not only output the location information and type information of each object, but also segment the image of the object from the first image to obtain the image of the object.

[0205] For example, such as Figure 12 As shown, the first image includes target object 1201 and target object 1202, wherein the regions where target object 1201 and target object 1202 are located do not overlap. Target detection is performed on the first image to obtain images 1203 and 1204, where image 1203 is an image of target object 1201 and image 1204 is an image of target object 1202.

[0206] like Figure 13 As shown, the first image includes target object 1205 and target object 1206. The areas where target object 1205 is located overlap with the areas where target object 1206 is located, and part of the content of target object 1206 occludes part of the content of target object 1205. Therefore, by performing target detection on the first image, we can obtain image 1207 and image 1208. Image 1207 is the image of target object 1205, and image 1208 is the image of target object 1206.

[0207] contrast Figure 12 , Figure 13 It can be seen that when the areas of two target objects overlap, the electronic device can determine which target object the content of the overlapping area belongs to, and segment the overlapping area into the image of the corresponding target object when segmenting the image.

[0208] S803 performs text recognition on the image of the target object to obtain the text within the target object.

[0209] Understandably, since text recognition is performed on an image of the target object, the recognized text is naturally located within the target object.

[0210] S804 performs entity recognition on the text within the target object to obtain entity information of the target text.

[0211] The entity information of the target text is used to determine the probability of the target text belonging to each of the various entity types. Since the text within the target object is located within the target object, there is no need to determine the positional relationship between the text and the target object; entity recognition can be performed directly on the text within the target object to obtain the entity information of the target text.

[0212] S805 determines service recommendation information based on the type information of the target object, the entity information of the target text, and the preset configuration information.

[0213] For a detailed description of S805, please refer to S508, which will not be repeated here.

[0214] S806 recommends services based on service recommendation information.

[0215] For a detailed description of S806, please refer to S509, which will not be repeated here.

[0216] Understandably, the method provided in this application embodiment can simplify the processing flow of electronic devices without determining the positional relationship between the target object and the target text, thereby improving the efficiency of recommendation services.

[0217] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.

[0218] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.

[0219] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0220] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0221] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0222] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0223] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0224] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A service recommendation method, characterized in that, include: Display the image to be processed, which includes the image and text of the target object; In response to the user's first operation on the image to be processed, the type information of the target object and the entity information of the target text are obtained. The target text is the text in the first text of the image to be processed that overlaps with the image area of ​​the target object. The first text is the text in the text of the image to be processed whose entity type is a preset entity type. Service recommendation information is determined based on the type information of the target object, the entity information of the target text, and preset configuration information. The type information of the target object is used to indicate the first probability that the target object is each of multiple types. The entity information of the target text is used to indicate the second probability that the target text is each of multiple entity types. The configuration information is used to indicate the third probability that text of each entity type appears in each type of object. The service recommendation information is used to indicate the matching degree between the target object of each type and the target text of each entity type. The target service is determined based on the service recommendation information. Display the target service.

2. The method according to claim 1, characterized in that, Obtaining entity information of the target text includes: Based on the location information of the first text and the location information of the target object's image, the positional relationship between the target object's image and the first text is determined. The positional relationship is used to indicate whether the area where the first text is located overlaps with the area where the target object's image is located. The entity information of the target text is obtained based on the positional relationship and the entity information of the first text. The entity information of the first text is used to indicate the second probability that the first text is each of a variety of entity types.

3. The method according to claim 2, characterized in that, The method further includes: Perform text recognition on the image to be processed to obtain the text and its location information; Entity recognition is performed on the text to obtain the first text and the entity information of the first text.

4. The method according to claim 2 or 3, characterized in that, Obtaining the type information of the target object includes: Target detection is performed on the image to be processed to obtain the type information of the target object and the position information of the target object in the image.

5. The method according to claim 2 or 3, characterized in that, The step of determining the positional relationship between the image of the target object and the first text based on the positional information of the first text and the positional information of the image of the target object includes: The candidate region and its location information are obtained by zooming in on the area where the first text is located. Based on the location information of the candidate region and the location information of the target object's image, the positional relationship between the target object's image and the first text is determined.

6. The method according to claim 1, characterized in that, Obtaining entity information of the target text includes: Perform text recognition on the image of the target object to obtain all the text within the target object; Entity recognition is performed on all text within the target object to obtain entity information of the target text.

7. The method according to claim 1 or 6, characterized in that, Obtaining the type information of the target object includes: Target detection is performed on the image to be processed to obtain the type information of the target object.

8. The method according to any one of claims 1-3 or 6, characterized in that, For target objects and target texts that do not overlap in their respective regions, the matching degree between the target object of type i and the target text of type j is positively correlated with the first probability that the target object is of type i, the second probability that the target text is of type j, and the third probability that text of type j appears under the object of type i.

9. The method according to claim 8, characterized in that, The type information of the target object, the entity information of the target text, the configuration information, and the service recommendation information satisfy the following: M i-j =P1 i *P2 j *P3 i-j Among them, P1 i P2 represents the first probability that the target object is of type i. j P3 represents the second probability that the target text is of the j-th entity type. i-j M represents the third probability that text of the j-th entity type will appear under the i-th type of object. i-j The matching degree between the target object of type i and the target text of type j is represented, where i ≤ N, j ≤ R, N is the number of the various types, and R is the number of the various entity types.

10. The method according to any one of claims 1-3 or 6, characterized in that, Determining the target service based on the service recommendation information includes: If the entity type of the target text matches the type of the target object based on the service recommendation information, the service corresponding to the entity type of the target text is determined to be the target service.

11. The method according to claim 10, characterized in that, The method further includes: If the maximum value in the matching degree is greater than or equal to a preset threshold, the entity type of the target text is determined to be the entity type corresponding to the maximum value, the type of the target object is determined to be the type corresponding to the maximum value, and the entity type of the target text matches the type of the target object.

12. The method according to claim 11, characterized in that, The method further includes: The target text is modified according to the type of the target object.

13. An electronic device, characterized in that, The electronic device includes: a memory and one or more processors; The memory is used to store computer program code, which includes computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, Includes computer instructions; When the computer instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Method and device for pushing services

    CN106326371A

  • Information recommendation method and electronic equipment

    CN115857737A