Service recommendation method, electronic equipment and computer readable storage medium
By combining the information of text and object, the electronic device can more accurately identify the information of interest to the user and provide related services based on this, solving the service-free problem caused by identification errors in the prior art and improving the user experience.
Patent Information
- Application Number
- CN202311673831.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-12-06
AI Technical Summary
When existing electronic devices identify the information of interest, due to the influence of light, occlusion, dirt, blur and other factors, the identified information is inconsistent with the real information, and the services recommended to users are not related to the real information of interest, which affects the user experience.
By combining the identified text and the object in which the text is located, a service recommendation method is provided to obtain the type information of the target object and the entity information of the target text, and determine the service recommendation information based on this information and preset configuration information, thereby improving the accuracy of the recommended service.
It improves the accuracy of service recommendations, enhances user experience, and ensures that the recommended services are related to real information of interest.
Smart Images

Figure CN120148040A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to a service recommendation method, an electronic device, and a computer-readable storage medium. Background Art
[0002] In daily life, information that users are interested in can be seen everywhere. For example, this information can exist in electronic images and paper documents. In order to improve the user experience, electronic devices such as mobile phones can identify information of interest in any form and recommend some services that match the information of interest to the user. For example, when the information of interest is identified as an express waybill number, services for tracking and querying the express are recommended to the user, etc.
[0003] However, since the information of interest is often affected by factors such as light, occlusion, dirt, and blur, the information of interest recognized by the electronic device is inconsistent with the real information of interest, which further leads to the services recommended by the electronic device having no association with the real information of interest, affecting the user experience. Summary of the Invention
[0004] Embodiments of this application provide a service recommendation method, an electronic device, and a computer-readable storage medium, which can recommend services to users by combining the recognized text and the object where the text is located, improve the accuracy of the recommended services, and improve the user experience.
[0005] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, this application provides a service recommendation method, including: displaying a to-be-processed picture, the to-be-processed picture including an image and text of a target object; in response to a first operation of a user on the to-be-processed picture, displaying a target service, the target service being associated with the type of the target object and the entity type of the target text of the to-be-processed picture, the target text being the text in the first text of the to-be-processed picture where the text area overlaps with the area where the image of the target object is located, and the first text being the text in the to-be-processed picture whose entity type is a preset entity type.
[0007] Based on the above solution, the electronic device can determine the target service by combining the entity type of the text and the type of the object where the text is located, increasing the dimension for reference when determining the target service, making the accuracy of the target service higher, and improving the user experience.
[0008] In an implementation provided in the first aspect, before displaying the target service, the method further includes: obtaining the type information of the target object and the entity information of the target text, where the type information of the target object is used to indicate the first probability of the target object being each of multiple types, and the entity information of the target text is used to indicate the second probability of the target text being each of multiple entity types; determining service recommendation information according to the type information of the target object, the entity information of the target text, and preset configuration information, where the configuration information is used to indicate the third probability of the text of each entity type appearing in each type of object, and the service recommendation information is used to indicate the matching degree between each type of target object and the target text of each entity type; determining the target service according to the service recommendation information.
[0009] It can be understood that when there are multiple target objects, the electronic device can obtain the type information of multiple target objects; when there are multiple target texts, the electronic device can obtain the entity information of multiple target texts.
[0010] In an implementation provided in the first aspect, obtaining the entity information of the target text includes: determining the positional relationship between the image of the target object and the first text according to the positional information of the first text and the positional information of the image of the target object, where the positional relationship is used to indicate whether the area where the first text is located overlaps with the area where the image of the target object is located; obtaining the entity information of the target text according to the positional relationship and the entity information of the first text, where the entity information of the first text is used to indicate the second probability of the first text being each of multiple entity types. Among them, determining the positional relationship between the image of the target object and the first text according to the positional information of the first text and the positional information of the image of the target object can accurately determine on which specific target object the first text is located, and then determine the target text.
[0011] In an implementation provided in the first aspect, the method further includes: performing text recognition on the picture to be processed to obtain the text and the positional information of the text; performing entity recognition on the text to obtain the first text in the text and the entity information of the first text.
[0012] In an implementation provided in the first aspect, obtaining the type information of the target object includes: performing target detection on the picture to be processed to obtain the type information of the target object and the positional information of the image of the target object.
[0013] In an implementation provided in the first aspect, determining the positional relationship between the image of the target object and the first text according to the positional information of the first text and the positional information of the image of the target object includes: magnifying the area where the first text is located to obtain a candidate area and the positional information of the candidate area; determining the positional relationship between the image of the target object and the first text according to the positional information of the candidate area and the positional information of the image of the target object.
[0014] Understandably, considering that the text itself occupies a relatively small proportion in the picture, in some cases, the actual area occupied by the text may exceed the text area determined by the position information of the text. Therefore, obtaining the candidate area by expanding the text area determined based on the position information of the text can more accurately determine the area occupied by the text, and the positional relationship between the image of the target object and the first text obtained in this way is also more accurate.
[0015] In an implementation provided in the first aspect, obtaining the entity information of the target text includes: performing text recognition on the image of the target object to obtain all the text within the target object; performing entity recognition on all the text within the target object to obtain the entity information of the target text. This method can simplify the processing flow of the electronic device without having to judge the positional relationship between the image of the target object and the target text, and can achieve the effect of improving the efficiency of the recommendation service.
[0016] In an implementation provided in the first aspect, obtaining the type information of the target object includes: performing target detection on the picture to be processed to obtain the type information of the target object.
[0017] In an implementation provided in the first aspect, for a target object and a target text whose regions do not overlap, the matching degree between the target object of the i-th type and the target text of the j-th entity type is positively correlated with the first probability that the target object is of the i-th type, the second probability that the target text is of the j-th entity type, and the third probability that the text of the j-th entity type appears under the object of the i-th type.
[0018] In an implementation provided in the first aspect, the type information of the target object, the entity information of the target text, the configuration information, and the service recommendation information satisfy:
[0019] M i-j = P1 i * P2 j * P3 i-j
[0020] Wherein, P1 i represents the first probability that the target object is of the i-th type, P2 j represents the second probability that the target text is of the j-th entity type, P3 i-j represents the third probability that the text of the j-th entity type appears under the object of the i-th type, and M i-j represents the matching degree between the target object of the i-th type and the target text of the j-th entity type, i ≤ N, j ≤ R, N is the number of multiple types, and R is the number of multiple entity types.
[0021] In an implementation provided in the first aspect, determining a target service according to service recommendation information includes: when it is determined that the entity type of the target text matches the type of the target object according to the service recommendation information, determining the service corresponding to the entity type of the target text as the target service.
[0022] In an implementation provided in the first aspect, the method further includes: if the maximum value in the matching degree is greater than or equal to a preset threshold, determining the entity type corresponding to the maximum value as the entity type of the target text, determining the type corresponding to the maximum value as the type of the target object, and determining that the entity type of the target text matches the type of the target object.
[0023] In an implementation provided in the first aspect, the method further includes: correcting the target text according to the type of the target object to improve the accuracy of the text recognition result.
[0024] In a second aspect, the present application provides an electronic device, which includes: a memory and one or more processors; wherein, the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the electronic device is caused to execute the method according to the first aspect and any one of its implementations.
[0025] In a third aspect, the present application provides a computer-readable storage medium, including computer instructions; when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to the first aspect and any one of its implementations.
[0026] Among them, for the technical effects brought by any one of the design manners in the second aspect to the third aspect, reference can be made to the technical effects brought by different design manners in the first aspect, which will not be elaborated here. Description of the Drawings
[0027] Figure 1 It is a schematic diagram of a scenario provided by an embodiment of the present application;
[0028] Figure 2 It is another schematic diagram of a scenario provided by an embodiment of the present application;
[0029] Figure 3 It is yet another schematic diagram of a scenario provided by an embodiment of the present application;
[0030] Figure 4 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;
[0031] Figure 5 It is a hierarchical architecture diagram of an electronic device provided by an embodiment of the present application;
[0032] Figure 6Interaction schematic diagram between modules provided by embodiments of the present application;
[0033] Figure 7 Flow schematic of the service recommendation method provided by embodiments of the present application Figure 1 ;
[0034] Figure 8 An example diagram of the first picture provided by embodiments of the present application;
[0035] Figure 9 Schematic diagram of the positional relationship between the image of the target object and the first text provided by embodiments of the present application;
[0036] Figure 10 Schematic diagram of the positional relationship between the first text and the candidate region provided by embodiments of the present application;
[0037] Figure 11 Flow schematic of the service recommendation method provided by embodiments of the present application Figure 2 ;
[0038] Figure 12 Schematic diagram of a segmented image provided by embodiments of the present application;
[0039] Figure 13 Another schematic diagram of a segmented image provided by embodiments of the present application. Detailed implementation manners
[0040] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be described. Among them, in the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "the", "above-mentioned", "this" and "this one" are also intended to include, for example, the expression form of "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0041] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First", "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features.
[0042] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0043] Currently, electronic devices such as mobile phones can identify information of interest to users in any form and recommend some services that match the information of interest to the users. Exemplarily, in the scenario where the user opens the gallery application, the mobile phone can display Figure 1 the interface 101 shown in (a) in, which is the interface of the gallery application. The interface 101 includes thumbnails of multiple pictures, for example, including the thumbnail 1011. In response to the user's operation on the thumbnail 1011, the mobile phone can display Figure 1 the interface 102 shown in (b) in. Among them, the interface 102 includes the picture 103 corresponding to the thumbnail 1011. There may be information of interest to the user in the picture 103, such as the text 1031 in the picture 103. Taking the picture 103 as a picture of a bank card as an example, the text 1031 of interest to the user may be the bank name, bank card number, etc. In addition, the interface 102 also includes an identification icon 104. In response to the user's operation on the identification icon 104, the mobile phone can identify the information of interest to the user, and the mobile phone can display service options that match the identification result.
[0044] Exemplarily, Figure 1In (c), service option 105 is shown when the mobile phone misidentifies the text 1031 (which is actually a bank card number) as a landline number (for example, 62366813). This service option 105 indicates that the number "62366813" can be called. Among them, in response to the user's operation of clicking on service option 105, the mobile phone can call the number "62366813". In response to the user's long - press operation on service option 105, the mobile phone can display Figure 1 the service window 106 shown in (d). The service window 106 includes options such as "Call 62366813", "Send Message", "Add to Contacts", etc. In response to the user's operation on the "Call 62366813" option, the mobile phone can call the number "62366813"; in response to the user's operation on the "Send Message" option, the mobile phone can display the text message window between the local device and the number "62366813"; in response to the user's operation on the "Add to Contacts" option, the mobile phone can add the number "62366813" to the contacts.
[0045] In other possible designs, the mobile phone may also misidentify the text 1031 as an express delivery number, and then can display service options such as querying express delivery and tracking express delivery.
[0046] It can be seen that the text recognized by the mobile phone is inconsistent with the real text, and the services recommended by the mobile phone to the user (such as the call service) have no relation to the real text. This is actually because currently the mobile phone only recommends services based on the recognized text, but the text recognized by image recognition is often affected by factors such as light, occlusion, dirt, and blur, making the recognized text inaccurate, further resulting in the services recommended by the electronic device to the user having no relation to the real text and affecting the user experience.
[0047] In view of this, the present application provides a service recommendation method, which can recommend services to the user by combining two dimensions: the text itself and the object where the text is located, and can improve the adaptation rate of the recommended services to the text and enhance the user experience.
[0048] The method provided in the embodiments of the present application can be applied to Figure 1 the scenario where the user opens the gallery application as shown, and can also be applied to scenarios such as screenshot scenarios, shooting scenarios, browsing scenarios, etc. Refer to Figures 2 - 3 , which is a schematic diagram of the application scenario provided by the present application. Exemplarily, in the screenshot scenario, as Figure 2As shown in (a) in the figure, the electronic device can display a screenshot 201, which may contain information of interest to the user. Taking the screenshot 201 as the details of the express delivery order, the text 2011 of interest to the user may be the order number, address, mobile phone number, etc. In response to the user pressing the screenshot 201 with two fingers, the electronic device can recognize the text 2011 in the screenshot 201 and display multiple service options, which are respectively matched with the entities recognized by the electronic device from the screenshot 201. For example, if the screenshot 201 includes the three entities of the express delivery order number, address, and phone number, then Figure 2 As shown in (b) in FIG. 8 , the electronic device may display three service options, which are “track the courier” 202 , “navigate to XX community” 203 , and “call 183XXXX1111” 204 .
[0049] For example, in a shooting scene, Figure 3 As shown in (a) of FIG. 3 , the electronic device can display a shooting preview interface 301. The captured image 302 displayed in the shooting preview interface 301 may contain information that the user is interested in. Taking the image 302 as an example of an ID card, the information that the user is interested in may be the ID card number, name, address, etc. In addition, the shooting preview interface 301 also includes an identification icon 303. In response to the user's operation on the identification icon 303, the electronic device can identify the information of interest in the image 302 and display multiple service options. Figure 3 As shown in (b) in FIG. 1 , the mobile phone may display a service option 304 of “Navigate to Xth Floor, No. X, XX Community, XX City, XX Province”.
[0050] The service recommendation method provided in the embodiments of the present application can be adapted to electronic devices, for example, the electronic devices may be mobile phones, tablet computers, televisions (also referred to as smart screens), desktop computers, laptop computers, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, wearable devices, vehicle-mounted devices, smart home devices and / or smart city devices, etc. The embodiments of the present application do not impose any special restrictions on the specific types of the electronic devices.
[0051] Figure 4 FIG. 1 shows a schematic diagram of the hardware structure of an electronic device. Figure 4As shown in the figure, the electronic device 200 may include: a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0052] Among them, the processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The processor 210 may be the nerve center and command center of the electronic device. The processor 210 may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0053] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0054] In the embodiments of the present application, the processor 210 may perform text recognition, entity recognition, object detection, etc. on pictures, and determine the services to be recommended to the user according to the corresponding results.
[0055] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0056] The external memory interface 220 may be used to connect to an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0057] The internal memory 221 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 210 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 221. For example, in the embodiments of the present application, the processor 210 may execute the instructions stored in the internal memory 221, and the internal memory 221 may include a program storage area and a data storage area.
[0058] Among them, the program storage area may store an operating system, applications required for at least one function (such as a service recommendation function, etc.). The data storage area may store data created during the use of the electronic device (such as configuration information, entity information of the target text, location information of the target text, entity information of the first text, location information of the target object, which will be specifically described later). In addition, the internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0059] The charging management module 240 is used to receive a charging input from a charger. Here, the charger can be a wireless charger or a wired charger. While charging the battery 242, the charging management module 240 can also supply power to the electronic device through the power management module 241.
[0060] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives inputs from the battery 242 and / or the charging management module 240 and supplies power to the processor 210, the internal memory 221, the external memory, the display screen 294, the camera 293, the wireless communication module 260, etc. In some embodiments, the power management module 241 and the charging management module 240 can also be provided in the same device.
[0061] The wireless communication function of the electronic device can be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, and the baseband processor, etc. In some embodiments, the antenna 1 of the electronic device is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, so that the electronic device can communicate with the network and other devices through wireless communication technologies.
[0062] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0063] The mobile communication module 250 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device. The mobile communication module 250 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves from the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation.
[0064] The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 250 can be provided in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 can be provided in the same device.
[0065] The wireless communication module 260 may provide wireless communication solutions applied to an electronic device, including WLAN (such as a (wireless fidelity, Wi-Fi) network), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc.
[0066] The wireless communication module 260 may be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 may also receive the signals to be sent from the processor 210, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 and radiate them out.
[0067] The electronic device may implement audio functions through the audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and the application processor, etc. For example, music playback, recording, etc.
[0068] The sensor module 280 may include sensors such as a pressure sensor, gyroscope sensor, barometric pressure sensor, magnetic sensor, acceleration sensor, distance sensor, proximity light sensor, fingerprint sensor, temperature sensor, touch sensor, ambient light sensor, and bone conduction sensor, etc. The electronic device may collect various data through the sensor module 280.
[0069] The electronic device implements display functions through the GPU, display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.
[0070] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. In the embodiments of the present application, the display screen 294 may be used to display the services determined by the processor 210 to be recommended to the user, as well as some related interfaces, such as the Figures 1 - 3 interfaces in the above.
[0071] The electronic device can implement the shooting function through the ISP, camera 293, video codec, GPU, display screen 294, application processor, etc. The ISP is used to process the data fed back by the camera 293. The camera 293 is used to shoot static images or videos. In some embodiments, the electronic device may include one or N cameras 293, where N is a positive integer greater than 1.
[0072] The keys 290 include a power-on key, volume keys, etc. The keys 290 can be mechanical keys or touch keys. The motor 291 can generate vibration prompts. The motor 291 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. The indicator 292 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 295 is used to connect the SIM card. The SIM card can be in contact with and separated from the electronic device by being inserted into or removed from the SIM card interface 295. The electronic device can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.
[0073] It can be understood that the interface connection relationship between the modules illustrated in this embodiment is only for illustrative purposes and does not constitute a structural limitation on the electronic device. In other embodiments, the electronic device may also include more or fewer modules than those provided in the above embodiments, and different interface connection methods or combinations of multiple interface connection methods may also be adopted between the various modules.
[0074] The software system of the above electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservices architecture, or cloud architecture. In the embodiments of the present invention, the layered architecture of the Android system is taken as an example to exemplarily illustrate the software structure of the electronic device.
[0075] As Figure 5 shown, the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through interfaces. In some embodiments, the Android system may include an application layer, an application framework layer, Android runtime and algorithm libraries, a hardware abstraction layer (HAL), and a kernel layer. It should be noted that the embodiments of the present application take the Android system as an example. In other operating systems (such as the IOS system, etc.), as long as the functions implemented by each functional module are similar to those of the embodiments of the present application, the solutions of the present application can also be implemented.
[0076] Among them, the application layer may include a series of applications. For example, the application layer may include applications such as a camera, a gallery, notes, document editing, navigation, service association, etc., which are not specifically limited herein. Among them, the service association application can provide functions such as text recognition and service recommendation for other applications (such as camera, gallery, etc.).
[0077] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, it may include a window manager, an activity manager, a resource manager, a notification manager, a camera framework, etc., and the embodiments of the present application do not make any restrictions on this.
[0078] The algorithm library may include multiple functional modules. For example, an object detection module, a text recognition module, an entity recognition module, a position relationship judgment module, etc. Among them, the object detection module is used to detect whether there is a preset type of object in the picture and determine the probability that the object is each of multiple types. The text recognition module is used to recognize the text in the picture and determine the position information of the text. The entity recognition module is used to judge whether the text is a preset entity type and determine the probability that the text is each of multiple entity types. The position relationship judgment module is used to judge the position relationship between the above-mentioned object and the text, and this position relationship is used to indicate whether there is an overlap between the area where the text is located and the area where the object is located.
[0079] Optionally, this algorithm library may also be referred to as the algorithm engine layer.
[0080] Android runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system. The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0081] The HAL layer is an encapsulation of the Linux kernel driver, providing an interface upward and shielding the implementation details of the underlying hardware.
[0082] The HAL layer may include a camera HAL, a sensor HAL, an audio HAL, etc.
[0083] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, an audio driver, a camera driver, etc.
[0084] The software modules involved in the service recommendation method provided in the embodiments of the present application and the interactions between the modules will be described below. As Figure 6 shown, the camera application in the application layer can send a request for detecting a picture to the service association application located in the same layer. Among them, the service association application can send a target detection request and a text recognition request to the target detection module and the text recognition module of the algorithm library respectively. The target detection module can perform target detection on the picture and obtain a target detection result, which includes the probabilities of various types of target objects. The text recognition module can perform text recognition on the picture to obtain a text recognition result, which includes the text and the position information of the text. The text recognition module can also input the text recognition result into the entity recognition module, and the entity recognition module can recognize the entities in the text and the probabilities of the text being various entity types. The target detection module can send the target detection result to the position relationship judgment module, and the entity recognition module can send the probabilities of the text being various entity types and the position information of the text to the position relationship judgment module. In this way, the position relationship judgment module can determine the position relationship between the text and the target object. Then, the position relationship judgment module can send the position relationship between the text and the target object, the probabilities of various types of target objects, and the probabilities of the text being various entity types to the service association application. The service association application can determine the matching degree between the text and the target object based on the above information and the preset configuration information, and then recommend a service based on the matching degree.
[0085] For the convenience of understanding, the service recommendation method provided in the embodiments of the present application will be specifically introduced below with reference to the accompanying drawings.
[0086] Refer to Figure 7 , which is a flowchart of the service recommendation method provided in the embodiments of the present application Figure 1 . As Figure 7 shown, the service recommendation method includes S501 to S509.
[0087] S501, obtain a first picture.
[0088] Among them, the first picture can be a picture pre-stored in the electronic device, such as Figure 1 the picture 103 shown in ; it can be a picture captured by the electronic device in real time, such as a picture captured by the electronic device and displayed in the shooting preview interface, such as Figure 3 the picture 302 shown in ; it can be a picture sent by other devices, such as a picture received through a chat application when the user chats with other users using the electronic device, etc. No specific limitation is made here. Optionally, the first picture can also be referred to as a picture to be processed.
[0089] In a possible design, on the one hand, the first picture may include an image of a target object. Herein, the target object refers to an object of a preset type, and the objects of the preset type include but are not limited to express delivery forms, bank cards, identity cards, business cards, books, screenshots, etc. On the other hand, the first picture may include text. The text may include Chinese characters, numbers, symbols, etc., and no specific limitation is made here.
[0090] Exemplarily, in the Figure 8 first picture as shown, the bank card 401 and the identity card 402 are both target objects. The first picture includes an image of the bank card 401 (the image area identified by the dotted line box outside the bank card 401 in the figure) and an image of the identity card 402 (the image area identified by the dotted line box outside the identity card 402 in the figure). In addition, the first picture also includes texts such as "XX Bank", "62236681XXXXXXXXXXX1", "ATM", "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", "Citizen ID Number: 5XXXXXXXXXXXXXXXXX", etc.
[0091] Also exemplarily, Figure 1 in the picture 103 as shown, the bank card is the target object, and the image of the area where the bank card is located is the image of the target object; Figure 2 in the screenshot 201 as shown, it is the target object, so the screenshot 201 itself is the image of the target object; Figure 3 in the picture 302 as shown, the identity card is the target object, and the image of the area where the identity card is located is the image of the target object.
[0092] S502. Perform target detection on the first picture to obtain the type information of the target object in the first picture and the position information of the image of the target object.
[0093] Herein, the type information of the target object is used to indicate the probability (also referred to as the first probability) of each of multiple types of the target object. In the embodiments of the present application, the multiple types include express delivery forms, bank cards, identity cards, business cards, books, screenshots, etc.
[0094] In a possible design, the electronic device may input the first picture into a target detection model. The target detection model is a model that has been pre-trained to convergence and can be used to detect target objects in pictures. The input of the target detection model is a picture, and the output of the target detection model is the position information of each target object in the picture and the type information of each target object. Optionally, Figure 5 the target detection module in
[0095] Exemplarily, for Figure 8For the first picture shown, the electronic device can obtain the type information of the bank card 401 and the type information of the identity card 402. Among them, the type information of the bank card 401 includes the probabilities indicating that the bank card 401 is respectively an express bill, a bank card, an identity card, a business card, and a book, and the type information of the identity card 402 includes the probabilities indicating that the identity card 402 is respectively an express bill, a bank card, an identity card, a business card, and a book.
[0096] The position information of the image of the target object is used to indicate the position of the image of the target object in the first picture.
[0097] In a possible design, the position information of the image of the target object may include the coordinates of the four vertices of the image of the target object. Among them, as Figure 8 shown, the coordinate system where the coordinates are located takes the upper left vertex of the first picture as the coordinate origin, the direction of the width of the first picture as the x-axis, and the direction of the height of the first picture as the y-axis. It should be noted that unless otherwise specified, the coordinates appearing in the following text are all referenced to the above coordinate system.
[0098] In a possible design, if the electronic device determines that there is no target object in the first picture after performing target detection on the first picture, the electronic device can end the process.
[0099] S503. Perform text recognition on the first picture to obtain the text in the first picture and the position information of the text.
[0100] In the embodiments of the present application, the electronic device can use an optical character recognition (OCR) algorithm to perform text recognition on the first picture to obtain the text in the first picture and the position information of the text. Optionally, Figure 5 the text recognition module in
[0101] It should be noted that the text can be all the text in the first picture or the text selected by the user in the first picture, and no specific limitation is made here.
[0102] Exemplarily, when the electronic device performs text recognition on Figure 8 the first picture shown, it can obtain texts such as "XX Bank", "62236681XXXXXXXXXXX1", "ATM", "Name Chen XX", "Address No. X, Building X, XX Community, XX District, XX City, XX Province", "Citizen ID Number 5XXXXXXXXXXXXXXXXX", etc., and the respective position information of these texts.
[0103] The position information of the text is used to indicate the position of the text in the first picture. In a possible design, the position information of the text may include the coordinates of the four vertices of the area where the text is located.
[0104] In the embodiments of the present application, the text in the first picture includes the first text. Among them, the first text is the text whose entity type in the text of the first picture is a preset entity type. The preset entity type is any one of express delivery number, mobile phone number, landline number, service number, address, bank card number, ID card number, email, address, website, name.
[0105] For example, texts such as "62236681XXXXXXXXXXX1", "Name Chen XX", "Address No. X, Building X, XX Community, XX District, XX City, XX Province", "Civil ID number 5XXXXXXXXXXXXXXXXX" are all the first text, and texts such as "XX Bank", "ATM" are not the first text.
[0106] It should also be noted that there is no strict execution order between S502 and S503 above. S502 can be executed first and then S503, S503 can be executed first and then S502, or S502 and S503 can be executed simultaneously. No specific limitation is made here.
[0107] S504, perform entity recognition on the text in the first picture to obtain the entity information of the first text.
[0108] Among them, the entity information of the first text is used to indicate the probability (also called the second probability) that the first text in the first picture is each of multiple entity types. The multiple entity types include express delivery number, mobile phone number, landline number, service number, address, bank card number, ID card number, email, address, website, name.
[0109] Exemplarily, Figure 8 In the shown first picture, the entity type of "62236681XXXXXXXXXXX1" is bank card number, the entity type of "Name Chen XX" is name, the entity type of "Address No. X, Building X, XX Community, XX District, XX City, XX Province" is address, the entity type of "Civil ID number 5XXXXXXXXXXXXXXXXX" is ID card number, and the entity types of texts such as "ATM", "XX Bank" do not belong to any of the above multiple entity types. Then the electronic device can determine that texts such as "62236681XXXXXXXXXXX1", "Name Chen XX", "Address No. X, Building X, XX Community, XX District, XX City, XX Province", "Civil ID number 5XXXXXXXXXXXXXXXXX" are all the first text, and texts such as "ATM" and "XX Bank" are not the first text.
[0110] In a possible design, the electronic device may input the text in the first picture obtained in S503 into a named entity recognition (NER) model. The NER model is a pre-trained model that can be used to extract entities in the text. The input of the NER model is text, and the output of the NER model includes whether each text is the first text, and the probability that the first text is each of multiple entity types. In the embodiments of the present application, Figure 5 The above NER model may be integrated in the entity recognition module shown.
[0111] Exemplarily, for Figure 8 the first picture shown, the electronic device may input texts such as "XX Bank", "62236681XXXXXXXXXXX1", "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", "Civil ID Number: 5XXXXXXXXXXXXXXXXX" into the NER model, and obtain the probabilities that texts such as "62236681XXXXXXXXXXX1", "Name: Chen XX", "Address: XX Village, XX County, XX City, XX Province", "Civil ID Number: 5XXXXXXXXXXXXXXXXX" are express delivery number, mobile phone number, landline number, service number, address, bank card number, ID number, email, address, website, and name respectively. Among them, the NER model may determine that the entity type of "XX Bank" does not belong to any of the above multiple entity types, and thus will not output the probabilities that "XX Bank" is express delivery number, mobile phone number, landline number, service number, address, bank card number, ID number, email, address, website, and name respectively.
[0112] In a possible design, if the electronic device determines that there is no first text in the text of the first picture after entity recognition, the electronic device may end the process.
[0113] S505, Screen and obtain the position information of the first text from the position information of the text.
[0114] It can be understood that, in the case where the position information of each text has been determined and which texts are the first text has been determined, the electronic device may directly screen and obtain the position information of the first text from the position information of the text.
[0115] S506, Obtain the positional relationship between the image of the target object and the first text according to the position information of the image of the target object and the position information of the first text.
[0116] In an embodiment of the present application, the positional relationship between a target object and a first text may include two types: far away and not far away. Among them, if the positional relationship between the first text and the target object is far away, it means that the first text is outside the image of the target object, or it can be understood that the area where the first text is located does not overlap with the area where the image of the target object is located. If the positional relationship between the first text and the target object is not far away, it means that part or all of the first text is located within the area where the image of the target object is located, or it can be understood that part or all of the area where the first text is located overlaps with the area where the image of the target object is located.
[0117] It should be noted that the positional relationship between the image of the target object and the first text described in the embodiment of the present application includes the positional relationship between any target object and any first text. For example, if there is one target object and one first text in the first picture, the electronic device can directly obtain the positional relationship between the target object and the first text. Another example is that if there are target object 1, target object 2, first text 1, first text 2, and first text 3 in the first picture, the positional relationship between the image of the target object and the first text includes the positional relationships between target object 1 and first text 1, first text 2, and first text 3 respectively, and the positional relationships between target object 2 and first text 1, first text 2, and first text 3 respectively. In this way, the electronic device can determine on which specific target object the first text is located.
[0118] In a possible design, for any first text and any target object, the electronic device can determine whether all four vertices of the area where the first text is located are outside the area where the image of the target object is located. If all four vertices of the area where the first text is located are outside the area where the image of the target object is located, it indicates that the area where the first text is located does not overlap with the area where the image of the target object is located, and thus the positional relationship between the first text and the image of the target object can be determined as far away; if at least one of the four vertices of the area where the first text is located is inside the area where the image of the target object is located, it indicates that the area where the first text is located partially overlaps with the area where the image of the target object is located, and thus the positional relationship between the first text and the image of the target object can be determined as not far away.
[0119] Refer to Figure 9, which shows a schematic diagram of the positional relationship between the image of the target object and the first text. Among them, all four vertices of the area where the first text 602 is located are within the area where the image of the target object 601 is located, so the positional relationship between the first text 602 and the image of the target object 601 is not far away. Two vertices of the area where the first text 603 is located are within the area where the image of the target object 601 is located, and two vertices are outside the area where the image of the target object 601 is located, so the positional relationship between the first text 603 and the image of the target object 601 is not far away. All four vertices of the area where the first text 604 is located are outside the area where the image of the target object 601 is located, so the positional relationship between the first text 604 and the image of the target object 601 is far away.
[0120] In an embodiment of the present application, the electronic device can determine whether the vertices of the first text are within the image of the target object according to the coordinates of each vertex of the first text and the coordinates of the four vertices of the image of the target object. Exemplarily, as Figure 9 shown, for any vertex P of the first text, the coordinates of this vertex P are (X, Y), and the position information of the target object includes the coordinates (X1, Y1), (X2, Y2), (X3, Y3) and (X4, Y4) of the four vertices of the target object. The electronic device can determine Xmax, Xmin, Ymax, and Ymin among the coordinates of the four vertices. Xmax is the maximum value of the abscissa among the four vertices, Xmin is the minimum value of the abscissa among the four vertices, Ymax is the maximum value of the ordinate among the four vertices, and Ymin is the minimum value of the ordinate among the four vertices. If Xmin ≤ X ≤ Xmax and Ymin ≤ Y ≤ Ymax, it means that the vertex P is within the target object; if X < Xmin or X > Xmax or Y < Ymin or Y > Ymax, it means that the vertex P is outside the target object.
[0121] Exemplarily, as Figure 9 shown, the position information of the target object includes the coordinates (X1, Y1) of the lower left vertex, the coordinates (X2, Y2) of the upper left vertex, the coordinates (X3, Y3) of the upper right vertex, and the coordinates (X4, Y4) of the lower right vertex. Among them, Xmin is X1, Xmax is X3, Ymin is Y2, and Ymax is Y4. Then, if the coordinates of the vertex P satisfy X1 ≤ X ≤ X3 and Y2 ≤ Y ≤ Y4, it is determined that the vertex P is within the target object; otherwise, it is determined that the vertex P is outside the target object.
[0122] In a possible design, the electronic device can determine whether the four vertices of the first text are within the target object through the following code
[0123] Public boolean contain(X,Y){
[0124] return Xmin < Xmax && Ymin < Ymax && X >= Xmin && X <= Xmax && Y >= Ymin && Y <= Ymax
[0125] }
[0126] In another possible design, the electronic device can first magnify the area where the first text is located to obtain the position information of the candidate area. Then, for this candidate area, the electronic device can determine whether all four vertices of the candidate area are located outside the target object. If all four vertices of the candidate area are located outside the target object, the positional relationship between the first text and the target object can be determined as being far away; if at least one of the four vertices of the candidate area is located inside the target object, the positional relationship between the first text and the target object can be determined as not being far away.
[0127] Exemplarily, as Figure 10 shown, based on the area 701 where the first text is located, the electronic device expands the area 701 where the first text is located by a at both ends of its height and expands the area 701 where the first text is located by b at both ends of its width to obtain the candidate area 702.
[0128] Taking the coordinates of the four vertices of the position information of the candidate area as (X5, Y5), (X6, Y6), (X7, Y7), and (X8, Y8) respectively, and the position information of the first text including the coordinates of the four vertices as (X9, Y9), (X10, Y10), (X11, Y11), and (X12, Y12) as an example, the position information of the candidate area and the position information of the first text satisfy:
[0129] (1) The square of the distance between the corresponding two vertices in the area where the first text is located and the candidate area is a fixed value:
[0130] (X9 - X5) 2 +(Y9 - Y5) 2 = a 2 + b 2 , (X10 - X6) 2 +(Y10 - Y6) 2 = a 2 + b 2 ,
[0131] (X11 - X7) 2 +(Y11 - Y7) 2 = a 2 + b 2 , (X12 - X8) 2 +(Y12 - Y8) 2 = a 2 + b 2
[0132] (2) The slopes of the corresponding two sides in the first text area and the candidate area are equal:
[0133] Where, when the rectangular frames where the first text area and the candidate area are located are not parallel to the coordinate system:
[0134] (Y6 - Y5) / (X6 - X5) = (Y10 - Y9) / (X10 - X9);
[0135] (Y7 - Y6) / (X7 - X6) = (Y11 - Y10) / (X11 - X10);
[0136] (Y8 - Y7) / (X8 - X7) = (Y12 - Y11) / (X12 - X11);
[0137] (Y8 - Y5) / (X8 - X5) = (Y12 - Y9) / (X12 - X9);
[0138] When the rectangular frames where the first text and the candidate area are located are parallel to the coordinate system:
[0139] X5 = X6, X10 = X9 = X5 + b, X7 = X8, X12 = X11 = X7 - b.
[0140] (3) The distances from the four vertices of the candidate area to the corresponding sides of the first text area are fixed values:
[0141] Taking the vertex (X6, Y6) of the candidate area as an example, the first text area includes side 1 formed by vertices (X10, Y10) and (X11, Y11), so the distance from the vertex (X6, Y6) to side 1 satisfies the following formula:
[0142]
[0143] Similarly, the first text area also includes side 2 formed by vertices (X11, Y11) and (X8, Y8), side 3 formed by vertices (X8, Y8) and (X9, Y9), and side 4 formed by vertices (X9, Y9) and (X10, Y10). Calculate the distance from the vertex (X7, Y7) of the candidate area to side 2 of the first text area, calculate the distance from the vertex (X8, Y8) of the candidate area to side 3 of the first text area, and calculate the distance from the vertex (X5, Y5) of the candidate area to side 4 of the first text area, and list the corresponding formulas.
[0144] According to the above three groups of arithmetic expressions, the electronic device can calculate a and b, and then determine the position information of the candidate area based on the position information of the first text and a, b. Then, the electronic device can further determine the positional relationship between the candidate area and the image of the target object based on the position information of the candidate area and the position information of the image of the target object, so as to determine that the positional relationship between the first text and the image of the target object is far away or not far away.
[0145] It can be understood that considering that the text itself occupies a relatively small proportion in the picture, in some cases, the actual area occupied by the text may exceed the text area determined by the position information of the text. Therefore, by expanding the text area determined based on the position information of the text to obtain the candidate area, the area occupied by the text can be determined more accurately.
[0146] Optionally, the electronic device can also expand the area where the image of the target object is located to obtain a candidate area, and then determine whether all four vertices of the area where the first text is located are outside the candidate area. If all four vertices of the area where the first text is located are outside the candidate area, it can be determined that the positional relationship between the first text and the target object is far away; if at least one vertex of the four vertices of the area where the first text is located is inside the candidate area, it can be determined that the positional relationship between the first text and the target object is not far away. Alternatively, the electronic device can expand the area where the first text is located and the area where the image of the target object is located respectively, and then determine the positional relationship between the first text and the target object based on the two expanded areas.
[0147] S507, obtain the entity information of the target text according to the positional relationship between the image of the target object and the first text and the entity information of the first text.
[0148] Among them, the target text is the text whose text area in the first text of the first picture overlaps with the area where the image of the target object is located.
[0149] The entity information of the target text is used to indicate the probability (which can also be called the second probability) of the target text being each of multiple entity types.
[0150] It can be understood that the electronic device has obtained the probability of the first text being each of multiple entity types (the entity information of the first text described above). Therefore, the electronic device can determine which target texts are specifically included in each image of the target object according to the positional relationship between the image of the target object and the first text, and then obtain the entity information of the target text.
[0151] Exemplarily, continue with Figure 8Taking the first picture shown as an example, it has been described above that texts such as "62236681XXXXXXXXXXX1", "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", and "Citizen ID Number: 5XXXXXXXXXXXXXXXXX" in this picture are the first texts. Among them, since the area where "62236681XXXXXXXXXXX1" is located overlaps with the area where the image of bank card 401 is located, the electronic device can determine that "62236681XXXXXXXXXXX1" is within the image of bank card 401. Also, since the areas where texts such as "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", and "Citizen ID Number: 5XXXXXXXXXXXXXXXXX" are located overlap with the area where the image of ID card 402 is located, the electronic device can determine that "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", and "Citizen ID Number: 5XXXXXXXXXXXXXXXXX" are within the image of ID card 402. Further, the electronic device can determine that "62236681XXXXXXXXXXX1", "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", and "Citizen ID Number: 5XXXXXXXXXXXXXXXXX" are all target texts. Thus, the entity information of the target texts includes the probabilities of "62236681XXXXXXXXXXX1", "Name: Chen XX", "Address: No. X, Building X, XX Community, XX District, XX City, XX Province", and "Citizen ID Number: 5XXXXXXXXXXXXXXXXX" respectively for each of the multiple entity types.
[0152] S508. Determine service recommendation information based on the type information of the target object, the entity information of the target text, and the preset configuration information.
[0153] Among them, the preset configuration information is used to indicate the probability (also referred to as the third probability) of a text of each entity type appearing under each type of object. It should be noted that each of the third probabilities in this configuration information can be obtained by statistics or experiments by developers, and in practical applications, it is the probability of various different entity types appearing under different types.
[0154] Exemplarily, the preset configuration information can be as shown in Table 1.
[0155] Table 1
[0156]
[0157]
[0158] Exemplarily, Table 1 indicates that the probability of the express delivery number appearing on the express delivery form is 95%, the probability of the mobile phone number appearing on the express delivery form is 80%, the probability of the bank card number appearing on the express delivery form is 5%, and the probability of the ID number appearing on the express delivery form is 5%; Table 1 also indicates that the probability of the express delivery number appearing on the bank card is 5%, the probability of the mobile phone number appearing on the bank card is 5%, the probability of the bank card number appearing on the bank card is 95%, and the probability of the ID number appearing on the bank card is 5%.
[0159] It should be noted that the types, entity types, and the third probability shown in Table 1 above are only illustrative. In other embodiments, the configuration information may further include more types and entity types than those in Table 1, and the third probability may also be adjusted.
[0160] The service recommendation information is used to indicate the matching degree between the target object of each type and the target text of each entity type.
[0161] In the embodiments of the present application, for the target object and the target text whose positional relationship is not far away, the matching degree between the target object of the i-th type and the target text of the j-th entity type is positively correlated with the following three parameters, and the following three parameters include: the first probability that the target object is of the i-th type, the second probability that the target text is of the j-th entity type, and the third probability that the text of the j-th entity type appears under the object of the i-th type. Wherein, i ≤ N, j ≤ R, N is the number of multiple types, and R is the number of multiple entity types.
[0162] In a possible design, the type information of the target object, the entity information of the target text, the configuration information, and the service recommendation information satisfy:
[0163] M i-j = P1 i * P2 j * P3 i-j
[0164] Wherein, P1 i represents the first probability that the target object is of the i-th type, P2 j represents the second probability that the target text is of the j-th entity type, P3 i-j represents the third probability that the text of the j-th entity type appears under the i-th type, and M i-j represents the matching degree between the target object of the i-th type and the target text of the j-th entity type.
[0165] The following first takes an example where a first picture includes a target object and a target text to illustrate the process of determining the service recommendation information according to the type information of the target object, the entity information of the target text, and the preset configuration information.
[0166] Exemplarily, the type information of the target object indicates that the probability of the target object being an express bill is 70%, the probability of being a bank card is 5%, the probability of being an ID card is 7%, the probability of being a business card is 13%, and the probability of being a book is 5%. The entity information of the target text indicates that the probability of the target text being an express bill number is 30%, the probability of being a mobile phone number is 6%, the probability of being a landline number is 6%, the probability of being a service number is 8%, the probability of being an address is 3%, the probability of being a bank card number is 40%, the probability of being an ID card number is 3%, the probability of being an email is 1%, the probability of being a website URL is 2%, and the probability of being a name is 1%.
[0167] According to Table 1, the probability of an express bill number appearing on an express bill is 95%. Then, when the target object is an express bill and the target text is an express bill number, the matching degree is 70% × 30% × 95% = 19.95%.
[0168] According to Table 1, the probability of a mobile phone number appearing on an express bill is 80%. Then, when the target object is an express bill and the target text is a mobile phone number, the matching degree is 70% × 6% × 80% = 3.36%.
[0169] According to Table 1, the probability of a landline number appearing on an express bill is 50%. Then, when the target object is an express bill and the target text is a landline number, the matching degree is 70% × 6% × 50% = 2.10%.
[0170] According to Table 1, the probability of a service number appearing on an express bill is 50%. Then, when the target object is an express bill and the target text is a service number, the matching degree is 70% × 8% × 50% = 2.80%.
[0171] According to Table 1, the probability of an address appearing on an express bill is 80%. Then, when the target object is an express bill and the target text is an address, the matching degree is 70% × 3% × 80% = 1.68%.
[0172] According to Table 1, the probability of a bank card number appearing on an express bill is 5%. Then, when the target object is an express bill and the target text is a bank card number, the matching degree is 70% × 40% × 5% = 1.40%.
[0173] According to Table 1, the probability of an ID card number appearing on an express bill is 5%. Then, when the target object is an express bill and the target text is an ID card number, the matching degree is 70% × 3% × 5% = 0.11%.
[0174] According to Table 1, the probability of an email appearing on an express bill is 50%. Then, when the target object is an express bill and the target text is an email, the matching degree is 70% × 1% × 50% = 0.35%.
[0175] According to Table 1, the probability of a website URL appearing on an express delivery form is 30%. Then, when the target object is an express delivery form and the target text is a website URL, the matching degree is 70% × 2% × 30% = 0.42%.
[0176] According to Table 1, the probability of a name appearing on an express delivery form is 80%. Then, when the target object is an express delivery form and the target text is a website URL, the matching degree is 70% × 1% × 80% = 0.6%.
[0177] And so on, the electronic device can obtain the service recommendation information as shown in Table 2.
[0178] Table 2
[0179]
[0180] When there are multiple target objects and multiple target texts in the first picture, the electronic device can determine the matching degree between any two target objects and target texts whose positional relationship is not far away according to the method described above.
[0181] For example, as Figure 8 shown, the first picture includes two target objects, a bank card 401 and an ID card 402. Among them, target text 1 (such as 62236681XXXXXXXXXXX1) is included in area 301, and target text 2 (such as "Chen XX"), target text 3 (such as "No. X, Building X, XX Community, XX District, XX City, XX Province") and target text 3 (such as "5XXXXXXXXXXXXXXXXX") are included in area 302.
[0182] Then the electronic device can obtain the type information of the bank card 401, the type information of the ID card 402, the entity information of target text 1, the entity information of target text 2, and the entity information of target text 3.
[0183] Then, since the positional relationship between target text 1 and the bank card 401 is not far away, and the positional relationships between target text 2 and target text 3 and the ID card 402 are not far away, the electronic device can determine the matching degree between the bank card 401 and target text 1 according to the type information of the bank card 401 and the entity information of target text 1, determine the matching degree between the ID card 402 and target text 2 according to the type information of the ID card 402 and the entity information of target text 2, and determine the matching degree between the ID card 402 and target text 3 according to the type information of the ID card 402 and the entity information of target text 3, so as to obtain the service recommendation information.
[0184] S509, perform service recommendation based on the service recommendation information.
[0185] Among them, the electronic device can determine the target service according to the service recommendation information and display the target service, so as to achieve service recommendation.
[0186] In an embodiment of the present application, the electronic device may determine the maximum value in the matching degree according to the service recommendation information. If the maximum value in the matching degree is greater than or equal to a preset threshold, the type corresponding to the maximum value may be determined as the type of the target object, the entity type corresponding to the maximum value may be determined as the entity type of the target text, and the target object and the target text match. Furthermore, the service corresponding to the entity type of the target text may be determined as the target service.
[0187] Among them, the preset threshold may be set according to actual needs and will not be specifically limited here.
[0188] Exemplarily, according to Table 2, the maximum value of the matching degree is the matching degree between the target object of the type of express bill and the target text of the entity type of express bill number, which is 19.95%. If the matching degree is greater than or the preset threshold, the type of the target object may be determined as the express bill, and the entity type of the target text may be determined as the express bill number. Therefore, the electronic device determines that the target service is the service related to the express bill number.
[0189] In a possible design, if the target text is an express bill number, the target service may be a service for tracking the express; if the target text is a mobile phone number, the target services may be services such as making a call and creating a new contact; if the target text is an address, the target service may be a navigation service; if the target text is a bank card number, the target service may be a transfer service; if the target text is a website address, the target service may be a service for searching the web page, etc., which will not be specifically limited here.
[0190] If the maximum value in the matching degree corresponding to the target text is less than the preset threshold, it indicates that the target object and the target text corresponding to the matching degree do not match. The electronic device may not recommend any service, or recommend general services such as copying and collecting, which will not be specifically limited here.
[0191] It should be noted that if multiple entity types of target texts can be determined according to the service recommendation information, the electronic device may recommend the services corresponding to the entity types of these multiple target texts respectively. Or, the electronic device may only recommend the services corresponding to one or several entity types of the target texts, which will not be specifically limited here.
[0192] Exemplarily, take Figure 2Taking the scene shown as an example, in response to the user's operation of pressing the screenshot 201 with two fingers, the electronic device can perform object detection and text recognition on the screenshot 201. Among them, the electronic device can determine that the screenshot 201 is the target object, and obtain the probability that the target object is each of the types such as express bill, bank card, ID card, business card, book, etc., and recognize the text in the screenshot 201, such as texts like "78744071XXXXXXX", "XX Community", "183XXXX1111", "Shipped", etc., as well as the position information of all texts. Then, the electronic device can perform entity recognition on texts such as "78744071XXXXXXX", "XX Community", "183XXXX1111", "Shipped", etc., and respectively obtain the probability that "78744071XXXXXXX", "XX Community", "183XXXX1111" are each of multiple entity types. Next, the electronic device can screen out the position information of the first text from the position information of all texts, and obtain the positional relationship between the first text and the target object based on the position information of the first text and the position information of the image of the target object. In this scene, the positional relationship between all the first texts and the target object is non-far away, so all the first texts are target texts. Then, the electronic device can determine, based on the probability that target texts such as "78744071XXXXXXX", "XX Community", "183XXXX1111" are each of multiple entity types, the probability that the target object is each of multiple types, and the preset configuration information: the matching degree between the target object of each type and the target text "78744071XXXXXXX" of each entity type, the matching degree between the target object of each type and the target text "XX Community" of each entity type, and the matching degree between the target object of each type and the target text "183XXXX1111" of each entity type, and further determine, based on the above matching degrees, that the entity type of the target text "78744071XXXXXXX" is the express bill number, determine that the entity type of the target text "XX Community" is the address, and determine that the entity type of the target text "183XXXX1111" is the mobile phone number. Therefore, the electronic device can recommend services corresponding to the express bill number, such as express tracking service; recommend services corresponding to the address, such as navigation service; and recommend services corresponding to the mobile phone number, such as call service.
[0193] Optionally, the form in which the electronic device recommends services can be to display corresponding service options on the interface. For example, as shown in (b) of Figure 2 service recommendations are made through these three service options: "Track Express" 202, "Navigate to XX Community", and "Call 183XXXX1111" 204.
[0194] In other embodiments, the electronic device may also recommend only any one or two of the express tracking service, navigation service, or call service, without specific limitation herein. In addition, the form of service recommendation may not be limited to Figures 1 - 3 the form shown, and may also be embodied in other forms.
[0195] It can be seen that when recommending services in this application, not only the text content itself is considered, but also the type of the object where the text is located is considered. The service corresponding to the text will be recommended only when the matching degree between the text and the object is higher than the threshold. This method is beneficial to improving the accuracy of the recommended services and enhancing the user experience.
[0196] In a possible design, the electronic device may also correct the target text whose positional relationship with the target object is not far away according to the type of the target object. This method is beneficial to improving the accuracy of the text recognition result.
[0197] For example, if the electronic device determines that the target text is "51XXXo1996XXXX11o5", its entity type is an ID number, and the type of the target object is an ID card. Since it is considered that the probability of the letter "o" appearing in the ID card scenario is extremely low, the electronic device may change the letter "o" to the number "0", thereby correcting the target text to "51XXX01996XXXX1105".
[0198] In the above embodiment, the electronic device performs target detection and text recognition on the first picture respectively, which enables the electronic device to execute the processes of target detection and text recognition in parallel, improving the processing efficiency.
[0199] In another possible design, an embodiment of the present application provides another service recommendation method, which may not need to judge the positional relationship between the target text and the image of the target object.
[0200] Refer to Figure 11 for the flowchart of the service recommendation method provided by the embodiment of the present application Figure 2 . As Figure 11 shown, this service recommendation method includes S801 to S806. It should be noted that some processes and principles of the embodiment of the present application and the embodiment shown above Figure 7 are similar, and will not be repeated herein. Refer to the previous description.
[0201] S801, obtain the first picture.
[0202] Regarding the description of the first picture, reference can be made to S501, which will not be repeated herein.
[0203] S802, perform target detection on the first picture to obtain the type information of the target object and the picture of the target object in the first picture.
[0204] In a possible design, the target detection model can not only output the location information of each target object and the type information of each target object, but also segment the image of the target object from the first picture to obtain the picture of the target object.
[0205] Exemplarily, as Figure 12 shown, the first picture includes a target object 1201 and a target object 1202, where the areas where the target object 1201 is located and the target 1202 is located do not overlap. Performing target detection on the first picture can obtain a picture 1203 and a picture 1204. The picture 1203 is the picture of the target object 1201, and the picture 1204 is the picture of the target object 1202.
[0206] As Figure 13 shown, the first picture includes a target object 1205 and a target object 1206, where the areas where the target object 1205 is located and the target 1206 is located overlap, and part of the content of the target object 1206 obscures part of the content of the target object 1205. Then, performing target detection on the first picture can obtain a picture 1207 and a picture 1208. The picture 1207 is the picture of the target object 1205, and the picture 1208 is the picture of the target object 1206.
[0207] Comparing Figure 12 and Figure 13 it can be seen that when the areas where two target objects are located overlap, the electronic device can determine which target object the content of the overlapping area specifically belongs to, and segment the overlapping area into the pictures of the corresponding target objects when segmenting the image.
[0208] S803. Perform text recognition on the picture of the target object to obtain the text inside the target object.
[0209] It can be understood that since text recognition is performed on the picture of the target object, the recognized text is naturally located inside the target object.
[0210] S804. Perform entity recognition on the text inside the target object to obtain the entity information of the target text.
[0211] Among them, the entity information of the target text is used for the probability that the target text is each of multiple entity types. Since the text inside the target object is all located inside the target object, there is no need to judge the position relationship between the text and the target object here, and directly performing entity recognition on the text inside the target object can obtain the entity information of the target text.
[0212] S805. Determine service recommendation information according to the type information of the target object, the entity information of the target text, and the preset configuration information.
[0213] Among them, the specific description of S805 can be referred to S508, and will not be elaborated here.
[0214] S806 is to perform service recommendation based on service recommendation information.
[0215] Among them, the specific description of S806 can be referred to S509, and will not be elaborated here.
[0216] It can be understood that the method provided by the embodiments of the present application can simplify the processing flow of the electronic device without judging the positional relationship between the target object and the target text, and can achieve the effect of improving the efficiency of the recommendation service.
[0217] The embodiments of the present application also provide a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above-mentioned electronic device, the electronic device is enabled to execute each function or step executed by the electronic device in the above-mentioned method embodiments.
[0218] The embodiments of the present application also provide a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to execute each function or step executed by the electronic device in the above-mentioned method embodiments.
[0219] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0220] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0221] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0222] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0223] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0224] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A service recommendation method, characterized in that, it includes: display a picture to be processed, the picture to be processed including an image and text of a target object; in response to a first operation of a user on the picture to be processed, display a target service, the target service being associated with the type of the target object and the entity type of the target text of the picture to be processed, the target text being the text in which the text area of the first text in the picture to be processed overlaps with the area where the image of the target object is located, and the first text being the text in the text of the picture to be processed whose entity type is a preset entity type.
2. The method according to claim 1, characterized in that, before displaying the target service, the method further includes: obtain the type information of the target object and the entity information of the target text, the type information of the target object being used to indicate the first probability that the target object is each of multiple types, and the entity information of the target text being used to indicate the second probability that the target text is each of multiple entity types; determine service recommendation information according to the type information of the target object, the entity information of the target text, and preset configuration information, the configuration information being used to indicate the third probability that text of each entity type appears in each type of object, and the service recommendation information being used to indicate the matching degree between each type of the target object and the target text of each entity type; determine the target service according to the service recommendation information.
3. The method according to claim 2, characterized in that, obtaining the entity information of the target text includes: determine the positional relationship between the image of the target object and the first text according to the position information of the first text and the position information of the image of the target object, the positional relationship being used to indicate whether the area where the first text is located overlaps with the area where the image of the target object is located; obtain the entity information of the target text according to the positional relationship and the entity information of the first text, the entity information of the first text being used to indicate the second probability that the first text is each of multiple entity types.
4. The method according to claim 3, characterized in that, the method further includes: perform text recognition on the picture to be processed to obtain the text and the position information of the text; perform entity recognition on the text to obtain the first text in the text and the entity information of the first text.
5. The method according to claim 3 or 4, characterized in that, obtaining the type information of the target object includes: perform target detection on the picture to be processed to obtain the type information of the target object and the position information of the image of the target object.
6. The method according to any one of claims 3-5, characterized in that, the determining the positional relationship between the image of the target object and the first text according to the position information of the first text and the position information of the image of the target object includes: enlarge the area where the first text is located to obtain a candidate area and the position information of the candidate area; Determine the positional relationship between the image of the target object and the first text according to the positional information of the candidate region and the positional information of the image of the target object.
7. The method according to claim 2, wherein, obtain the entity information of the target text, including: perform text recognition on the image of the target object to obtain all the text within the target object; perform entity recognition on all the text within the target object to obtain the entity information of the target text.
8. The method according to claim 2 or 7, wherein, obtain the type information of the target object, including: perform target detection on the picture to be processed to obtain the type information of the target object.
9. The method according to any one of claims 1-8, wherein, For target objects and target texts whose regions do not overlap, the matching degree between the i-th type of target object and the j-th entity type of target text is positively correlated with the first probability that the target object is of the i-th type, the second probability that the target text is of the j-th entity type, and the third probability that the text of the j-th entity type appears under the i-th type of object.
10. The method according to claim 9, wherein, The type information of the target object, the entity information of the target text, the configuration information, and the service recommendation information satisfy: M i-j = P1 i * P2 j * P3 i-j Among them, P1 i represents the first probability that the target object is of the i-th type, P2 j represents the second probability that the target text is of the j-th entity type, P3 i-j represents the third probability that the text of the j-th entity type appears under the object of the i-th type, M i-j represents the matching degree between the target object of the i-th type and the target text of the j-th entity type, i ≤ N, j ≤ R, N is the number of the multiple types, and R is the number of the multiple entity types.
11. The method according to any one of claims 2-8, wherein, The determining the target service according to the service recommendation information includes: When it is determined that the entity type of the target text matches the type of the target object according to the service recommendation information, determine the service corresponding to the entity type of the target text as the target service.
12. The method according to claim 11, wherein, The method further includes: If the maximum value in the matching degree is greater than or equal to a preset threshold, determine that the entity type of the target text is the entity type corresponding to the maximum value, determine that the type of the target object is the type corresponding to the maximum value, and determine that the entity type of the target text matches the type of the target object.
13. The method according to claim 12, wherein, The method further includes: Correct the target text according to the type of the target object.
14. An electronic device, wherein, The electronic device includes: a memory and one or more processors; wherein, the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-13.
15. A computer-readable storage medium, wherein, including computer instructions; When the computer instructions run on an electronic device, the electronic device executes the method according to any one of claims 1-13.
Citation Information
Patent Citations
Method and device for pushing services
CN106326371A
Text information recommendation method and device, server and storage medium
CN114398549A
Application program recommendation method and electronic equipment
CN115033153A
Information recommendation method and electronic equipment
CN115857737A
Application recommendation method
CN116126197A