Schedule information generation method, electronic equipment, storage medium and product
By calling the text and intent recognition module on electronic devices to filter out schedule-irrelevant information and combining image categories to identify schedule information, the problem of low accuracy of schedule information in traditional methods is solved, achieving higher recognition accuracy and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional methods for recognizing calendar information from images do not produce accurate calendar information.
By calling the text recognition and intent recognition modules through the calendar application on the electronic device, text information unrelated to the schedule is filtered out, and the schedule information is recognized by combining the schedule information recognition model with image categories.
It improves the accuracy and efficiency of schedule information recognition, reduces interference, and enhances the user experience.
Smart Images

Figure CN121963230A_ABST
Abstract
Description
Methods for generating schedule information, electronic devices, storage media and products Technical Field
[0001] This application relates to the field of terminal technology, and in particular to methods for generating schedule information, electronic devices, storage media, and products. Background Technology
[0002] In daily life, users have many daily tasks to handle, and electronic devices can be set to remind users of these tasks.
[0003] To improve convenience, electronic devices can recognize images related to schedules to generate corresponding schedule information. For example, an electronic device can recognize hotel bookings and generate schedule information corresponding to those bookings. However, traditional methods often produce schedule information with low accuracy when recognizing schedule information from images.
[0004] Therefore, it is necessary to propose a method to address the problem of low accuracy in recognizing schedule information in scenarios based on image recognition. Summary of the Invention
[0005] This application provides a method for generating schedule information, an electronic device, a storage medium, and a product that can improve the accuracy of schedule information recognition.
[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0007] Firstly, a method for generating schedule information is provided, applied to an electronic device. A target image is received via a calendar application on the electronic device. The calendar application then calls a text recognition module to identify the image category corresponding to the target image and to identify the first text information within the target image. Next, the calendar application calls an intent recognition module to filter the first text information based on the text features corresponding to the image category, obtaining the remaining second text information. Finally, the calendar application sends the second text information and the image category to a schedule information recognition model, obtaining the schedule information corresponding to the target image output by the schedule information recognition model.
[0008] In the above scheme, the electronic device filters text information to prevent irrelevant text from interfering with the recognition process of the schedule information recognition model. Furthermore, the schedule information recognition model identifies schedule information based on the remaining filtered text, resulting in more accurate information. Additionally, image categories can be input as auxiliary information into the schedule information recognition model to further improve the accuracy of the identified information.
[0009] In one possible implementation of the first aspect, the text features corresponding to the image category include the positional features and / or style features of the text in the image under that image category. The process involves calling the intent recognition module via a calendar application on an electronic device to filter the first text information based on the text features corresponding to the image category, obtaining the remaining second text information after filtering. This includes: calling the intent recognition module via a calendar application on an electronic device to filter out text in the first text information that matches the positional features and / or style features, thus obtaining the second text information.
[0010] In the above scheme, text that interferes with the schedule information is filtered by the style features and / or position features of the text. This can prevent these texts that are not related to the actual schedule information from causing significant interference to the recognition of subsequent schedule information and improve the accuracy of subsequent schedule information recognition.
[0011] In another possible implementation of the first aspect, the text features corresponding to the image category include: keyword information corresponding to the image category. The intent recognition module is invoked through a calendar application on the electronic device to filter the first text information based on the text features corresponding to the image category, obtaining the remaining second text information after filtering. This includes: invoking the intent recognition module through the calendar application on the electronic device to filter out text in the first text information that matches the keyword information, obtaining the remaining second text information after filtering.
[0012] In the above scheme, the electronic device can also filter out interfering text that is unrelated to the actual schedule information by preset keywords, so as to prevent these texts that are unrelated to the actual schedule information from causing great interference to the recognition of subsequent schedule information and improve the accuracy of subsequent schedule information recognition.
[0013] In another possible implementation of the first aspect, the text recognition module is invoked through a calendar application on the electronic device to identify the image category corresponding to the target image. This includes: using the calendar application on the electronic device to invoke the text recognition module to perform preliminary identification of the target image's category, obtaining an initial category corresponding to the target image. If the initial category is the first target category, the electronic device further identifies the subcategory corresponding to the first target category as the image category corresponding to the target image.
[0014] In the above scheme, the initial category is merely a simple division of the image to be recognized, providing limited reference information for subsequent recognition processes. Advanced recognition of the subcategories corresponding to the images to be recognized by electronic devices allows for more targeted and effective filtering of text information from different image categories. Furthermore, the refinement of image categories provides more reference information to the calendar information recognition model, enabling it to obtain more accurate calendar information.
[0015] In another possible implementation of the first aspect, the image category includes any one of chat info images, notification images, order images, or non-order images.
[0016] In the above solution, refining the image categories allows electronic devices to more effectively filter text information from different image categories.
[0017] In another possible implementation of the first aspect, the schedule information recognition model includes at least one sub-model. Sending second text information and an image category to the schedule information recognition model via a calendar application on an electronic device to obtain the schedule information corresponding to the target image output by the schedule information recognition model includes: sending the second text information and an image category to the schedule information recognition model via a calendar application on an electronic device, causing the schedule information recognition model to determine a target sub-model corresponding to the image category, and analyzing the second text information through the target sub-model to obtain the schedule information corresponding to the target image.
[0018] In the above scheme, determining a more suitable sub-model based on image category can more accurately identify schedule information and improve the accuracy of schedule information recognition.
[0019] In another possible implementation of the first aspect, the schedule information generation method also includes a word segmentation process. The word segmentation process can be executed in parallel with the schedule information generation process. The steps of the word segmentation process include: calling the word segmentation interface of the intent recognition module through a calendar application on an electronic device to segment the second text information and obtain the word segmentation result corresponding to the target image.
[0020] In the above scheme, the electronic device executes the word segmentation process and the schedule information generation process in parallel, which can improve the processing efficiency and save the computing power of the electronic device.
[0021] In another possible implementation of the first aspect, after the electronic device obtains the word segmentation result corresponding to the target image, it further includes: displaying a first interface through a calendar application on the electronic device, wherein the first interface includes schedule information and schedule title, and the schedule title is generated based on the word segmentation result.
[0022] In the above solution, displaying schedule information and schedule titles on the interface can help remind users of their to-do items and improve the user experience.
[0023] In another possible implementation of the first aspect, after the first interface is displayed via a calendar application, the calendar application on the electronic device responds to the editing operation of the event title by displaying the word segmentation result corresponding to the target image. The calendar application on the electronic device then responds to the selection operation of the target word in the word segmentation result and modifies the event title based on the target word.
[0024] In the above solution, the schedule title can be modified by editing, thus enabling customization of the schedule title and improving the user experience.
[0025] In another possible implementation of the first aspect, after the text recognition module identifies the image category corresponding to the target image through the calendar application on the electronic device, if the image category is the second target category, the calendar application on the electronic device sends the target image and image category to the schedule information recognition model to obtain the schedule information corresponding to the target image. If the image category is not the second target category, the electronic device performs the step of recognizing the first text information in the target image.
[0026] In the above solution, the electronic device does not need to convert the image into text information through the text extraction plugin in the text recognition module, nor does it need to filter the text information. The electronic device can directly send the image to the schedule information recognition model on the cloud server, which can effectively improve the efficiency of schedule information generation and reduce the schedule information generation time.
[0027] Secondly, this application provides an electronic device comprising: a memory and one or more processors, the memory being coupled to the processors. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device performs the method described in the first aspect and any of its possible implementations.
[0028] Thirdly, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in the first aspect and any possible implementation thereof.
[0029] Fourthly, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to perform the method as described in the first aspect and any possible implementation thereof. The computer may be an electronic device as described in the third aspect and any possible implementation thereof.
[0030] Understandably, the beneficial effects achieved by the electronic device of the second aspect, the computer-readable storage medium of the third aspect, and the computer program product of the fourth aspect provided above can be referred to with reference to the beneficial effects of the method of the first aspect and any possible implementation thereof, which will not be repeated here. Attached Figure Description
[0031] Figure 1 is a schematic diagram of a ticket order screenshot provided in an embodiment of this application;
[0032] Figure 2 is a schematic diagram of the interface interaction of a schedule information generation method provided in an embodiment of this application;
[0033] Figure 3 is a schematic diagram of the interface interaction of another schedule information generation method provided in the embodiment of this application;
[0034] Figure 4 is a schematic diagram of the interface after generating schedule information according to an embodiment of this application;
[0035] Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0036] Figure 6 is a schematic diagram of the software system architecture of an electronic device provided in an embodiment of this application;
[0037] Figure 7 is a timing diagram of a schedule information generation method provided in an embodiment of this application;
[0038] Figure 8 is a schematic diagram of a first category of images provided in an embodiment of this application;
[0039] Figure 9 is a schematic diagram of another type of first-category image provided in an embodiment of this application;
[0040] Figure 10 is a schematic diagram of a hotel order screenshot provided in an embodiment of this application;
[0041] Figure 11 is a timing diagram of another schedule information generation method provided in an embodiment of this application;
[0042] Figure 12 is a timing diagram of another schedule information generation method provided in an embodiment of this application;
[0043] Figure 13 is a timing diagram of another schedule information generation method provided in an embodiment of this application;
[0044] Figure 14 is a timing diagram of another schedule information generation method provided in an embodiment of this application;
[0045] Figure 15 is a flowchart illustrating a method for generating schedule information according to an embodiment of this application;
[0046] Figure 16 is a schematic diagram of an interface for editing schedule information provided in an embodiment of this application. Detailed Implementation
[0047] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0048] With the development of technology, electronic devices can generate corresponding schedule information for users by recognizing images related to their schedules. These images can be captured by a camera or screenshots of schedule information displayed on the electronic device. Specifically, the schedule application (or calendar application) on the electronic device receives the image, recognizes and analyzes it to obtain the corresponding schedule information, and then creates a new schedule based on that information. However, the accuracy of schedule information obtained through image recognition is currently not high.
[0049] To address this, this application provides a method for generating schedule information, improving the accuracy of schedule information obtained based on image recognition. Specifically, in the embodiments of this application, after the electronic device receives an image (hereinafter referred to as the "image to be recognized") containing schedule information, it can classify the image to obtain the corresponding category, denoted as the image category. For example, the electronic device can determine the image category of the image to be recognized through a text recognition module, which can also be called an Optical Character Recognition (OCR) module. The electronic device can use the image category as auxiliary information and input it into the schedule information recognition model, enabling the electronic device to obtain the schedule information corresponding to the image to be recognized through the schedule information recognition model.
[0050] The schedule information recognition model can be set up in an electronic device or on a cloud server. If the schedule information recognition model is set up in an electronic device, it can be called a device-side large model; if the schedule information recognition model is set up on a cloud server, it can be called a cloud-side large model.
[0051] Understandably, in this embodiment, the electronic device inputs the image category as auxiliary information into the schedule information recognition model, so that the schedule information recognition model can refer to more information when recognizing schedule information in the image, thereby improving the recognition accuracy.
[0052] In some embodiments, after obtaining the image category, the electronic device can identify the schedule information in the image using either method one or method two.
[0053] Method 1: Electronic devices send images and image categories to the schedule information recognition model.
[0054] Specifically, in Method 1: After classifying the image, the electronic device can send the image and its category to the schedule information recognition model, allowing the model to identify the schedule information within the image based on the category. For example, the electronic device can send the image and its category to the schedule information recognition model via an intent recognition module. This intent recognition module can also be called a Natural Language Understanding (NLU) module.
[0055] Method 2: Electronic devices can also send the text information and image category recognized from the image to the schedule information recognition model.
[0056] Specifically, in Method Two, after classifying the image, the electronic device identifies the text within it to obtain the corresponding text information. The electronic device can directly send the identified text information (also known as the initially identified text information) and the image category to the schedule information recognition model. Alternatively, the electronic device can filter out irrelevant text from the identified text information based on the image category and send the remaining text information and image category to the schedule information recognition model. The schedule information recognition model can then identify the schedule information corresponding to the image based on the received text information, or based on the received text information and image category.
[0057] It is understandable that electronic devices recognize text in images to obtain text information, and the schedule information recognition model then uses the recognized text information to obtain schedule information, which can make the schedule information obtained by the schedule information recognition model more accurate.
[0058] In some examples, in Method 2, the electronic device can send text information from the image (such as text information initially identified from the image or text information remaining after filtering the initially identified text information) and the image category to the schedule information recognition model via the intent recognition module.
[0059] In some examples, the step of filtering recognized text information based on image category may include: the electronic device determining the text features in the image based on the image category, and filtering out text information unrelated to the schedule based on these text features. The electronic device can then send the remaining filtered text information and image category to the schedule information recognition model.
[0060] Understandably, electronic devices filter out irrelevant text information to prevent it from interfering with the calendar information recognition model's process. Since the model identifies calendar information based on the remaining filtered text, the final calendar information obtained is more accurate. Furthermore, image categories can be used as supplementary input to the calendar information recognition model to further improve its accuracy.
[0061] It should be understood that the methods in the embodiments of this application are not limited to the methods one or two described above, and there are other methods to identify schedule information in images based on image categories. For example, electronic devices can send the remaining text information after filtering to the schedule information recognition model separately.
[0062] The method described in this application can be applied to various scenarios, such as identifying schedules from orders, identifying schedules from chat content, or identifying schedules from notification information. Identifying schedules from orders includes identifying schedule information from any type of order, such as train ticket orders, hotel orders, flight orders, and performance orders.
[0063] For ease of understanding, the following text uses the scenario of identifying schedule information from a screenshot of a train ticket order as an example, and illustrates the application scenario of the schedule information generation method of this application embodiment with reference to Figures 1 to 4.
[0064] Please refer to Figure 1. After a user purchases a ticket using a travel-related application, the electronic device responds to the user's screenshot operation by taking a screenshot of the current ticket purchase interface, generating a ticket order screenshot 101, and displaying the ticket order screenshot 101 in the album interface.
[0065] Electronic devices can send screenshots of ticket orders to calendar applications using any of the following three methods.
[0066] Method 1: Send a screenshot of your ticket order to the calendar application via the share function.
[0067] In some embodiments, as shown in Figure 1 above, a "Share" button 102 is provided below the ticket order screenshot. In response to the "Share" button 102 being touched, the electronic device displays a sharing interface 201, as shown in Figure 2. The electronic device then sends the ticket order screenshot to the calendar application in response to the calendar application icon 202 on the sharing interface 201 being touched.
[0068] Method 2: Send a screenshot of the ticket order to the calendar application via the target application.
[0069] In some embodiments, cross-application data transfer between different applications can be achieved through a target application on the electronic device. Specifically, in response to the target content in the interface being dragged to the edge area of the screen (in which case multiple application icons can be displayed on the screen), the electronic device then sends the target content to the target application in response to the target content being released from the target application's application icon.
[0070] In one example, the electronic device responds to the current display interface 301 where a screenshot of a ticket order 101 is dragged to the edge of the display screen. In this case, as shown in Figure 3, an icon 202 of a calendar application can be displayed at the edge of the screen of the electronic device. When the electronic device then responds to the ticket order screenshot 101 being released from the calendar application icon 202, the ticket order screenshot will be sent to the calendar application.
[0071] Method 3: Import a screenshot of your train ticket order into the calendar application interface.
[0072] In some embodiments, the electronic device displays the interface of a calendar application, which includes a schedule creation entry. In response to a schedule creation trigger, the electronic device imports a screenshot of the ticket order into the calendar application through the schedule creation entry, thereby enabling the calendar application to receive the ticket order screenshot.
[0073] After the calendar application on the electronic device receives a screenshot of the ticket order, the electronic device then recognizes the screenshot, determines the schedule information corresponding to the ticket order screenshot, and displays the schedule information on the screen, as shown in Figure 4. For example, the electronic device can display schedule information 401 and the corresponding schedule title 402, etc., where schedule information may include: travel time, travel location, etc.
[0074] In some embodiments, sending a screenshot of a train ticket order to a calendar application on an electronic device is not limited to the three methods described above, and no restrictions are imposed here.
[0075] For example, the aforementioned electronic device may be a mobile phone, tablet computer, smart remote control, wearable device (such as smart bracelet, smartwatch, or smart glasses), PDA, augmented reality (AR) / virtual reality (VR) device. Alternatively, the electronic device 500 may also be a portable multimedia player (PMP), media player, or other types of electronic device. This application embodiment does not impose any limitations on the specific type of electronic device.
[0076] Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0077] Electronic device 500 may include processor 510, external memory interface 520, internal memory 521, universal serial bus (USB) interface 530, charging management module 540, power management module 541, battery 542, antenna 1, antenna 2, mobile communication module 550, wireless communication module 560, audio module 570, speaker 570A, receiver 570B, microphone 570C, headphone jack 570D, sensor module 580, button 590, motor 591, indicator 592, camera 593, display screen 594, and subscriber identification module (SIM) card interface 595, etc. The sensor module 580 may include a pressure sensor 580A, a gyroscope sensor 580B, a barometric pressure sensor 580C, a magnetic sensor 580D, an accelerometer sensor 580E, a distance sensor 580F, a proximity light sensor 580G, a fingerprint sensor 580H, a temperature sensor 580J, a touch sensor 580K, an ambient light sensor 580L, a bone conduction sensor 580M, etc.
[0078] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 500. In other embodiments of this application, the electronic device 500 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0079] Processor 510 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.
[0080] The controller can be the nerve center and command center of the electronic device 500. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.
[0081] The processor 510 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can store instructions or data that the processor 510 has just used or that are used repeatedly. If the processor 510 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.
[0082] In some embodiments, the processor 510 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0083] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 510 may include multiple I2C buses. The processor 510 can couple to the touch sensor 580K, charger, flash, camera 593, etc., through different I2C bus interfaces. For example, the processor 510 can couple to the touch sensor 580K through the I2C interface, enabling the processor 510 and the touch sensor 580K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 500.
[0084] The MIPI interface can be used to connect the processor 510 to peripheral devices such as the display screen 594 and the camera 593. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 510 and the camera 593 communicate via the CSI interface to enable the electronic device 500 to capture images. The processor 510 and the display screen 594 communicate via the DSI interface to enable the electronic device 500 to display images.
[0085] The GPIO interface is configurable via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 510 to a camera 593, a display screen 594, a wireless communication module 560, an audio module 570, a sensor module 580, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0086] Electronic device 500 implements display functions through a GPU, a display screen 594, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 594 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. Processor 510 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0087] Display screen 594 is used to display images, videos, etc. Display screen 594 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 500 may include one or N displays 594, where N is a positive integer greater than 1.
[0088] Electronic device 500 can achieve shooting function through ISP, camera 593, video codec, GPU, display 594 and application processor.
[0089] Camera 593 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 500 may include one or N cameras 593, where N is a positive integer greater than 1.
[0090] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 500 is selecting a frequency, the DSP is used to perform Fourier transforms on the frequency energy.
[0091] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn. NPUs can enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0092] Internal memory 521 can be used to store executable program code, including instructions. Processor 510 executes various functional applications and data processing of electronic device 500 by running the instructions stored in internal memory 521. Internal memory 521 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 500 (such as audio data, phonebook, etc.). Furthermore, internal memory 521 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0093] Touch sensor 580K, also known as a "touch panel," can be located on display screen 594. The touch sensor 580K and display screen 594 together form a touchscreen, also known as a "touchscreen." Touch sensor 580K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 594. In other embodiments, touch sensor 580K may also be located on the surface of electronic device 500, in a different position than display screen 594.
[0094] Figure 6 is a schematic diagram of the software system architecture of an electronic device provided in an embodiment of this application.
[0095] The software system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of an electronic device.
[0096] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into three layers, from top to bottom: the application layer, the application framework layer, and the driver layer.
[0097] The application layer can include a series of application packages.
[0098] As shown in Figure 6, the application package may include: a calendar application, a text recognition module, an intent recognition module, etc. It should be understood that it may also include: camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS, and other applications, not all of which are shown in the figure.
[0099] The calendar application receives the image to be recognized, obtains the schedule information from the image, and displays it on the screen.
[0100] The text recognition module is used to identify the text in the image to be recognized and obtain the text information.
[0101] The intent recognition module is used to filter out text information that may interfere with the subsequent process of recognizing schedule information.
[0102] In some embodiments, if the schedule information recognition model is set on a cloud server, the intent recognition module can also send a request to the schedule information recognition model, enabling the schedule information recognition model to recognize the remaining text information after filtering carried in the request, and obtain the schedule information. In other embodiments, if the schedule information recognition model is set on an electronic device, the application layer also includes a schedule information recognition model for recognizing the remaining text information after filtering by the intent recognition module, or recognizing the image to be recognized, to obtain the schedule information.
[0103] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes a set of predefined functions.
[0104] As shown in Figure 6, the application framework layer may also include: window manager, content provider, view system, phone manager, resource manager, notification manager, etc., which are not all shown in the figure.
[0105] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0106] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0107] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0108] The driver layer, also known as the kernel layer, includes display drivers, etc.
[0109] Display drivers enable the operating system to access and control the display screen, ensuring that images and video content are displayed correctly and efficiently on the screen.
[0110] The process of obtaining schedule information in Method 2 of this application is described below, taking into account the hardware and system structures described above:
[0111] 1. After receiving the image to be recognized, the calendar application sends the image to the text recognition module. The text recognition module recognizes the image and obtains the image category and text information corresponding to the image. The text recognition module then sends the recognized image category and text information back to the calendar application.
[0112] 2. The calendar application sends text information to the intent recognition module, which filters the text information to obtain the remaining text information. The intent recognition module then sends the remaining text information to the calendar application.
[0113] 3. The calendar application assembles and typeset the remaining filtered text information and sends the assembled and typeset text information to the intent recognition module.
[0114] 4. The intent recognition module then sends the assembled and formatted text information to the schedule information recognition model, which recognizes the text information to obtain the schedule information.
[0115] The process of obtaining schedule information in Method 1 of this application will be described below, taking into account the hardware and system architecture described above:
[0116] 1. After receiving the image to be recognized, the calendar application sends the image to the text recognition module. The text recognition module recognizes the image and obtains the image category corresponding to the image.
[0117] 2. The calendar application uses the intent recognition module to send the image category corresponding to the image to be recognized and the image to be recognized to the schedule information recognition model. The schedule information recognition model recognizes the image to be recognized and obtains the schedule information.
[0118] The method for generating schedule information provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0119] As discussed above, regardless of method one or method two, electronic devices can identify the image category corresponding to the image to be identified, and use this image category to assist the schedule information recognition model in identifying schedule information from the image. The following sections will provide a more detailed introduction to (i) image category identification and (ii) the schedule information recognition model.
[0120] (I) Introduction to identifying image categories.
[0121] In some embodiments, the specific process of identifying image categories includes: multiple image categories are pre-set in the electronic device. The electronic device can determine the image category corresponding to the image of the schedule information to be identified from the pre-set multiple image categories, thereby achieving the purpose of classifying the "image to be identified". For example, the electronic device can use a text recognition model to determine the image category corresponding to the image to be identified from the pre-set multiple image categories.
[0122] In some examples, the electronic device can use the image category determined from a set of pre-defined image categories as the final image category corresponding to the image to be identified.
[0123] In other examples, the electronic device may also use an image category determined from a pre-set set of multiple image categories as the initial category. After performing initial classification of the image to be identified, the electronic device can then perform further classification. Specifically, after determining the initial category corresponding to the image, the electronic device can further perform category recognition on the image to identify the sub-category corresponding to the image under that initial category, which will be the final image category corresponding to the image to be identified.
[0124] In some examples, after identifying the initial category corresponding to the image, the electronic device can determine whether the initial category is a preset target category 1 (or the first target category). If so, the electronic device can further perform category recognition on the image to be identified, to identify the sub-category corresponding to the image under the initial category, which is the final image category corresponding to the image to be identified. For example, target category 1 could be an order image as mentioned below. It should be understood that target category 1 could also be other image categories.
[0125] In some embodiments, the pre-defined multiple image categories may include three image categories: a first category, a second category, and a third category. The first category is chat info images, the second category is information notification images, and the third category is any category other than the first and second categories. A chat info image is an image that includes chat information or chat content, such as a screenshot of a chat interface or a screenshot of a chat history viewing interface. An information notification image is an image that contains communication information, such as a screenshot of a notification card used for notification purposes.
[0126] In some examples, if the image to be identified is initially identified as belonging to the third category mentioned above, the electronic device can further perform category identification on the image to be identified in order to determine the subcategory corresponding to the third category.
[0127] Understandably, the initial categorization is merely a simple division of the images to be recognized, providing limited reference information for subsequent recognition processes. Advanced recognition of subcategories for images allows electronic devices to more effectively filter text information from different image categories. Furthermore, the refinement of image categories provides more reference information to the calendar information recognition model, leading to more accurate calendar information.
[0128] Specifically, the third category can include: order-related images and non-order-related images. Order-related images can include: train ticket order images (such as orders generated by train ticket apps), hotel order images (such as orders generated by hotel apps), flight ticket order images (such as orders generated by flight ticket apps), performance order images (such as orders generated by performance ticketing apps), SMS order images, and order images generated by other applications (such as orders generated by shopping apps), etc.
[0129] In some examples, Method 1 can be executed if the image category to be identified is a preset target category 2 (or a second target category). For example, the preset target category 2 could be an order image. Specifically, if the image category to be identified is an order image, the electronic device can execute Method 1. If the image category to be identified is a category other than an order image, the electronic device can execute Method 2.
[0130] (II) Introduction to the schedule information recognition model.
[0131] As described above, in this embodiment of the application, a schedule information recognition model is used to identify schedule information from images. This schedule information recognition model can be a single, unified model or it can include multiple sub-models. There is no limitation on this, as long as it has the function of identifying schedule information from images.
[0132] In some embodiments, the schedule information recognition model is a single, unified model capable of identifying schedule information in images of any type or category. Therefore, electronic devices can input the image to be recognized, or the text information within the image, along with the identified image category, into the schedule information recognition model. This model can then use the image category to assist in identifying the schedule information in the image, resulting in more accurate schedule information and improving the accuracy of schedule information recognition.
[0133] In other embodiments, the schedule information recognition model is a multimodal large model that may include multiple sub-models. Each sub-model corresponds to an image category, and different sub-models obtain schedule information with different accuracies for different image categories. That is, each sub-model is responsible for recognizing schedule information from images of a particular image category. After inputting the image category as auxiliary information into the schedule information recognition model, the model can determine the sub-model corresponding to that image category. This sub-model is suitable for extracting schedule information from the image to be recognized. Furthermore, the schedule information can be processed using this sub-model to identify the image or text information within the image, thus obtaining the schedule information corresponding to the image. Therefore, determining a more suitable sub-model based on the image category can more accurately identify schedule information and improve the accuracy of schedule information recognition.
[0134] For example, if there are three image categories, denoted as Category 1, Category 2, and Category 3, the schedule information recognition model can include three sub-models. Sub-model 1 can accurately extract schedule information from images belonging to Category 1, sub-model 2 can accurately extract schedule information from images belonging to Category 2, and sub-model 3 can accurately extract schedule information from images belonging to Category 3. Assuming the image to be recognized is image A, and the electronic device can identify image A as belonging to Category 1, then when performing schedule information recognition processing on image A, in addition to inputting image A or the text information within image A into the schedule information recognition model, the electronic device can also input Category 1 into the schedule information recognition model. This allows the schedule information recognition model to call sub-model 1 corresponding to Category 1 to analyze image A or the text information within image A and obtain the schedule information.
[0135] The following text uses a scenario where a schedule information recognition model is set up on a cloud server, and the electronic device includes a calendar application, an intent recognition module, and a text recognition module, and schedule information is obtained through method two, to illustrate the schedule information generation method in detail.
[0136] In some embodiments, the electronic device first receives the target image through a calendar application. The electronic device then uses the calendar application to invoke a text recognition module to identify the image category corresponding to the target image, as well as the first text information within the target image. The electronic device then uses the calendar application to invoke an intent recognition module to filter the first text information based on the text features corresponding to the image category, obtaining the remaining second text information. The electronic device then uses the calendar application to send the second text information and the image category to a schedule information recognition model, obtaining the schedule information corresponding to the target image output by the schedule information recognition model.
[0137] The specific process of obtaining schedule information through method two can be shown in Figure 7, including the following steps:
[0138] Step 701: The calendar application sends an image to the text recognition module.
[0139] Step 702: The text recognition module returns the image category and the text information recognized from the image to the calendar application.
[0140] In some embodiments, after receiving an image from a calendar application, the text recognition module in the electronic device can classify the image to obtain its corresponding image category. The text recognition module divides the input image into smaller regions and then recognizes the text in each region. The text recognition module then sends the recognized text information and image category to the calendar application.
[0141] As an example, as shown in Figure 8, after receiving Figure 8, the text recognition module recognizes Figure 8 as an image of the first category. After recognizing the text in Figure 8, the text recognition module can obtain the following text information: "Okay", "Everyone eat first, I'll go from Zhongcheng", "Want to ride for a while", "Waiting for you", "November 24, 2022, 11:32", "Czy withdrew a message", "1124 interface people have a meal (9)", "This location, we'll start eating at 12 o'clock", "Jiangchuan Dim Sum (Shangmeilin Store)" and so on.
[0142] Step 703: The calendar application sends text information to the intent recognition module.
[0143] Step 704: The intent recognition module filters text information based on the text features corresponding to the image category, obtains the remaining text information after filtering, and sends the remaining text information after filtering to the calendar application.
[0144] In some embodiments, the intent recognition module can filter text information based on the text features corresponding to the image category to selectively filter out invalid text information that is unrelated to the schedule information, prevent invalid text information from interfering with the schedule information recognition model, enhance the recognition capability of the schedule information recognition model, and enable the schedule information recognition model to more accurately recognize the schedule information.
[0145] Step 705: The calendar application sends the filtered text information and image categories to the schedule information recognition model through the intent recognition module.
[0146] Step 706: The schedule information recognition model returns schedule information to the calendar application through the intent recognition module.
[0147] In some embodiments, in addition to filtering text information based on the text features corresponding to the image category, text information that does not contain location or time can also be filtered out.
[0148] As shown in Figure 8, after the text recognition module recognizes the text information in Figure 8, it can obtain the following text information: "Okay", "Everyone eat first, I'll go from Zhongcheng", "Want to ride for a while", "Waiting for you", "November 24, 2022, 11:32", "Czy withdrew a message", "1124 interface people have a meal (9)", "This location, we'll start eating at 12 o'clock", "Jiangchuan Dim Sum (Shangmeilin Store)". Among them, the text information such as "Okay", "Want to ride for a while", and "Waiting for you" do not contain time and location. In this case, the intent recognition module can filter the text information such as "Okay", "Want to ride for a while", and "Waiting for you".
[0149] In some embodiments, the text features corresponding to an image category may include the positional and / or style features of the text within the image. In other embodiments, the text features corresponding to an image category may also include keywords within the image.
[0150] The following section will describe in detail how to filter text information unrelated to the schedule based on the text features in the image, specifically including filtering method 1 and filtering method 2.
[0151] Filtering Method 1: Filter text information that is irrelevant to the schedule based on the position and / or style characteristics of the text in the image.
[0152] In some embodiments, the intent recognition module is invoked through a calendar application to filter out text in the first text information that matches the location features and / or style features, thereby obtaining the second text information.
[0153] In filtering method 1, the images may contain text unrelated to the actual schedule information, such as time and / or location text. This unrelated text can significantly interfere with the subsequent recognition of schedule information, affecting the accuracy of the recognition. For example, images in the first category (i.e., chat screenshots) may include time and / or location text that interferes with the schedule information, such as prompts for the chat message sending time or location text in location links.
[0154] The following section describes text filtering processing using the example of filtering time and / or location text that is irrelevant to the actual schedule information.
[0155] For ease of description, time and / or location text unrelated to the actual schedule information are designated as "interference text," while time and / or location text related to the actual schedule information are designated as "schedule-related text." Through in-depth research, the inventors of this application have discovered that the style and / or positional features of interference text differ from those of schedule-related text. Therefore, the inventors of this application propose a solution to filter time and / or location text that interferes with the schedule information based on its style and / or positional features. For example, an intent recognition module capable of filtering text based on style and / or positional features has been developed. This intent recognition module filters time and / or location text that interferes with the schedule information based on its style and / or positional features, preventing these times and locations unrelated to the actual schedule information from significantly interfering with the recognition of subsequent schedule information.
[0156] In some embodiments, style features may include: font size, slant angle, and height of the text line containing the text.
[0157] The inventors of this application have further investigated and found that the style and / or positional features of interfering text in images may differ for different image categories. Therefore, after determining the image category of the image to be identified, the electronic device can determine the style and positional features corresponding to that image category through an intent recognition module. It should be understood that the style and positional features corresponding to that image category refer to the style and positional features of the text in images within that image category. The electronic device can filter out text that matches the style and positional features corresponding to that image category from the text information recognized in the image (i.e., the initially recognized text information), thereby filtering out interfering text. The remaining text information after filtering is the text more relevant to the schedule.
[0158] Taking the first category of chat message images as an example, if the notification text indicating the chat message sending time is centered in the image and its font size is smaller than a preset font size, the intent recognition module will filter out the notification text indicating the chat message sending time. Similarly, if the image also includes location links, the text recognition module will identify the location links and obtain multiple location texts. Among these, at least one location text has a font size smaller than a preset font size, and at least one location text has an angle greater than a preset angle threshold between its horizontal and vertical orientations. In this case, the intent recognition module will filter out these location texts.
[0159] For easier understanding, a more detailed explanation is provided in conjunction with Figures 8 and 9.
[0160] As shown in Figure 8, the electronic device can obtain text information such as "November 24, 2022, 11:32", "Czy withdrew a message", "1124 Interface People Gathering (9)", "This location, start eating at 12 o'clock", "Jiangchuan Dim Sum (Shangmeilin Store)", "You Balu", "Fragrant Area", and "Okay". Among them, since the text "November 24, 2022, 11:32" and "Czy withdrew a message" are located in the center of the image, the intent recognition module filters these texts. Similarly, the tilt angle between the arrangement direction and the horizontal direction of the text "You Balu" and "Fragrant Area" is greater than the preset angle threshold, so the intent recognition module filters these texts. Similarly, the font size of the text "Lu Xinyu" and "Czy" is smaller than the preset font size, so the intent recognition module filters these texts. It can be understood that after the intent recognition module filters the text information corresponding to Figure 8, the text information includes: "1124 Interface People Gathering (9)", "This location, start eating at 12 o'clock", and "Jiangchuan Dim Sum (Shangmeilin Store)".
[0161] As shown in Figure 9, the electronic device can receive text information such as "January 31st, 22:31", "February 8th, 16:27", "February 8th, 16:37", "Are you free tonight?", "I get off work at 6:00", "I can get off work at 5:20", and "Yuhua Living Room". Since the text "January 31st, 22:31", "February 8th, 16:27", and "February 8th, 16:37" are centered in the image, the intent recognition module filters these texts. Understandably, after filtering the text information corresponding to Figure 9, the intent recognition module retains the following text: "Are you free tonight?", "I get off work at 6:00", "I can get off work at 5:20", and "Yuhua Living Room".
[0162] Taking the second category of information notification images as an example, in these images, the notification card is positioned in the center of the screen, with the current time displayed above it. If the current time text is centered and its font size is larger than a preset font size, the intent recognition module can filter out the current time text.
[0163] As can be seen, the interfering text corresponding to the first category of images and the second category of images have different style features and / or position features in the images. Therefore, electronic devices can set different style features and position features according to different categories of images to filter interfering text.
[0164] Filtering Method 2: Filter text information unrelated to the schedule based on keywords in the image.
[0165] In some embodiments, the calendar application invokes the intent recognition module to filter out text in the first text information that matches the keyword information, resulting in the remaining second text information. It is understood that in this embodiment, the electronic device can also filter out interfering text unrelated to the actual schedule information using preset keywords.
[0166] In some embodiments, the keywords corresponding to images of different image categories can also be different. Specifically, the distracting text may differ for different image categories. Therefore, different keywords can be preset to filter distracting text in images of different image categories, so that the remaining text information after filtering is more relevant to the schedule.
[0167] In some embodiments, for orders generated by train ticket booking apps, hotel booking apps, flight booking apps, performance booking apps, SMS orders, and other apps in the third category, text information can be filtered using keywords. For example, for images of hotel orders, keywords could be: "map / navigation," "taxi," "contact the merchant," "ask the hotel," "expected arrival," etc.
[0168] As an example, Figure 10(a) shows a screenshot of a hotel order. The text recognition module recognizes the text on the hotel order screenshot and obtains multiple text information, as shown in Figure 10(b). These text information include: text information 1001, text information 1002, text information 1003, text information 1004, text information 1005, text information 1006, text information 1007, text information 1008, text information 1009, and text information 1010. The intent recognition module can first filter out text information unrelated to the location and time information, i.e., filter text information 1001, text information 1002, text information 1008, etc. Secondly, since the font size of text information 1006 is smaller than the preset font size, the intent recognition module can also filter out text information 1006 and other text information using filtering method 1. Furthermore, since the text information 1004 contains the keyword "ask about hotels", the text information 1004 is filtered to prevent it from interfering with the generation of subsequent schedule information.
[0169] The specific execution steps of Method 2 will be explained in detail below with reference to Figure 11.
[0170] Step 1101: The calendar application is initialized.
[0171] Step 1102: The calendar application sends the image to the image classification plugin in the text recognition module.
[0172] In some embodiments, a calendar application in an electronic device sends an image as input data to a plugin or program in the text recognition module for image classification. For example, taking a bitmap (BMP) format image as an example, after the calendar application sends the BMP image to the image classification plugin, the image classification plugin analyzes and processes the BMP image to determine its category. The image classification plugin is a trained visual model.
[0173] Step 1103: The text recognition module returns the image category to the calendar application.
[0174] Step 1104: The calendar application sends the image to the text extraction plugin in the text recognition module.
[0175] In some embodiments, the text extraction plugin in the text recognition module receives an image, recognizes the text in the image, and divides the text into blocks according to the distribution area of the text to obtain at least one text information.
[0176] Step 1105: The text recognition module returns text information to the calendar application.
[0177] In some embodiments, during the process of the text recognition module recognizing text information from an image, the text information can also be sorted to improve the readability and ease of use of text recognition, as well as the accuracy of text recognition.
[0178] Step 1106: The calendar application calls the interface corresponding to the intent recognition module to filter the text information.
[0179] In some embodiments, the calendar application calls the entity recognition interface of the intent recognition module, and uses the text features corresponding to the image category to filter text information through the entity recognition interface.
[0180] Step 1107: The intent recognition module returns the filtered text information to the calendar application.
[0181] Step 1108: The calendar application assembles and typesets the filtered text information.
[0182] In some embodiments, calendar applications may assemble and format textual information to facilitate the accurate identification of schedule information from textual information by the schedule information recognition model, thereby making the schedule information identified by the schedule information recognition model more accurate.
[0183] As an example, the text information recognized by the text recognition module is shown in Figure 10(b). In Figure 10(b), "after 14:00 on September 1st (Friday), 2 nights, before 12:00 on September 3rd (Sunday)" is divided into different text information, namely, text information 1005, text information 1009, and text information 1010. However, in this case, it is not conducive to the schedule information recognition model to recognize accurate schedule information. Therefore, the text information can be assembled and formatted so that "after 14:00 on September 1st (Friday), 2 nights, before 12:00 on September 3rd (Sunday)" is in the same text information, that is, as shown in Figure 10(c), "after 14:00 on September 1st (Friday), 2 nights, before 12:00 on September 3rd (Sunday)" is in text information 1011.
[0184] One method for assembling and arranging text information is to establish a coordinate system on the image, using the origin of the coordinate system as a reference point, and then assemble and arrange the text information from left to right and from top to bottom. For example, as shown in Figure 10(c), text information 1001, text information 1002, and text information 1008 can be assembled into text information 1012 through this method.
[0185] Step 1109: The calendar application sends the assembled and formatted text information and image categories to the intent recognition module.
[0186] Step 1110: The intent recognition module sends the preprocessed and assembled text information and image categories to the schedule information recognition model.
[0187] In some embodiments, preprocessing may involve adding line breaks and spaces to the text information to facilitate the extraction of schedule information by the subsequent schedule information recognition model. After the intent recognition module preprocesses the text information, the text information and image category are sent to the schedule information recognition model.
[0188] Step 1111: The schedule information recognition model obtains schedule information by preprocessing, assembling, and formatting text information and image categories.
[0189] Step 1112: The schedule information recognition model sends the schedule information to the intent recognition module.
[0190] Step 1113: The intent recognition module sends the schedule information to the calendar application.
[0191] In some embodiments, the schedule information sent by the intent recognition module to the calendar application may be in JSON format.
[0192] Step 1114: The calendar application displays schedule-related information on the user interface.
[0193] The following is a detailed explanation of obtaining schedule information through method one:
[0194] In some embodiments, if the image category is the second target category, the electronic device sends the target image and image category to the schedule information recognition model via a calendar application to obtain the schedule information corresponding to the target image. If the image category is not the second target category, the electronic device performs the step of recognizing the first text information in the target image.
[0195] Specifically, as shown in Figure 12, the process of obtaining schedule information through Method 1 is as follows:
[0196] Step 1201: The calendar application is initialized.
[0197] Step 1202: The calendar application sends the image to the image classification plugin in the text recognition module.
[0198] Step 1203: The text recognition module returns the image category to the calendar application.
[0199] Step 1204: The calendar application sends the image and image category to the intent recognition module.
[0200] Specifically, the calendar application sends images and image categories to the intent recognition module's software development kit (SDK).
[0201] Step 1205: The intent recognition module sends the image and image category to the schedule information recognition model.
[0202] Step 1206: The schedule information recognition model identifies schedule information through images and image categories.
[0203] Step 1207: The schedule information recognition model sends the schedule information corresponding to the image to the intent recognition module.
[0204] Step 1208: The intent recognition module sends the schedule information to the calendar application.
[0205] Step 1209: The calendar application displays schedule-related information on the user interface.
[0206] Understandably, in Method 1, electronic devices do not need to convert images into text information through the text extraction plugin in the text recognition module, nor do they need to filter the text information. Electronic devices can directly send images to the schedule information recognition model on the cloud server. Therefore, it can effectively improve the efficiency of schedule information generation and reduce the schedule information generation time.
[0207] In some embodiments, as shown in Figure 4 above, the electronic device can display not only schedule information but also schedule titles. The schedule titles can be obtained through word segmentation, which involves dividing the text information into multiple words.
[0208] In some embodiments, the word segmentation interface of the intent recognition module can be called through the calendar application to segment the second text information and obtain the word segmentation result corresponding to the target image.
[0209] Specifically, electronic devices can also send the text recognized from images to an intent recognition module to obtain word segmentation. The word segmentation process and the schedule information processing process can be executed in parallel. Alternatively, the electronic device can obtain the word segmentation first and then the schedule information. Or, the electronic device can obtain the schedule information first and then the word segmentation.
[0210] Specifically, the following example illustrates this by taking the parallel execution of the word segmentation process and the schedule information processing process of the electronic device.
[0211] In some embodiments, regardless of whether the electronic device obtains the schedule information through method one or method two, the process of obtaining word segmentation and the process of obtaining schedule information can be executed in parallel.
[0212] Specifically, as shown in Figure 13, based on Method 2, steps 1301-1303 are also included:
[0213] Step 1301: The calendar application sends the assembled and formatted text information to the intent recognition module.
[0214] Step 1302: The intent recognition module calls the word segmentation interface to segment the text information.
[0215] Step 1303: The intent recognition module sends the string array corresponding to the word segmentation to the calendar application.
[0216] During the execution of steps 1109-1113, the electronic device also simultaneously executes steps 1301-1303. After obtaining the schedule information in step 1113 and the string array corresponding to the word segmentation in step 1303, step 1114 is executed, allowing the schedule information and the schedule title generated by word segmentation to be displayed on the screen. The electronic device can select at least one word from the string array corresponding to the word segmentation as the schedule title.
[0217] The following example, using a screenshot of a chat log, illustrates how to obtain calendar information using method two:
[0218] After receiving a chat screenshot, the calendar application first categorizes the screenshot using the image classification plugin in its text recognition module, assigning it to the first category. Then, the text extraction plugin in the text recognition module identifies the text within the screenshot, obtaining multiple text entries. The calendar application then calls the intent recognition module's interface to filter the text, obtaining the filtered text. Finally, the calendar application assembles and typesets the filtered text, resulting in the final, formatted text.
[0219] The calendar application then inputs the assembled and formatted text information and image categories into the intent recognition module. The intent recognition module then uses a schedule information recognition model to extract the schedule information corresponding to the images. The intent recognition module then returns the schedule information to the calendar application. During the process of obtaining the schedule information, the intent recognition module can also call a word segmentation interface to segment the text information and transmit the segmentation results to the calendar application via a string array.
[0220] In some other embodiments, as shown in FIG14, based on Method 1, steps 1401-1405 are also included.
[0221] Step 1401: The calendar application sends the image to the text extraction plugin in the text recognition module.
[0222] Step 1402: The text recognition module returns text information to the calendar application.
[0223] Step 1403: The calendar application sends text information to the intent recognition module.
[0224] Step 1404: The intent recognition module calls the word segmentation interface to segment the text information into words.
[0225] Step 1405: The intent recognition module sends the string array corresponding to the word segmentation to the calendar application.
[0226] During the execution of steps 1401-1405, the electronic device also simultaneously executes steps 1204-1208. After obtaining the schedule information in step 1208 and the string array corresponding to the word segmentation in step 1405, step 1209 is executed, allowing the schedule information and the schedule title generated by word segmentation to be displayed on the screen. The electronic device can select at least one word from the string array corresponding to the word segmentation as the schedule title.
[0227] The following example, using a screenshot of a hotel booking, illustrates how to obtain schedule information using Method 1:
[0228] After receiving a screenshot of a hotel order, the calendar application on the electronic device first classifies the screenshot using the image classification plugin in the text recognition module, and the classification result is the third category.
[0229] The calendar application then inputs the image and image category into the intent recognition module. The intent recognition module then uses a schedule information recognition model to extract the schedule information corresponding to the image. The schedule information recognition model then returns the schedule information to the calendar application via the intent recognition module.
[0230] During the process of obtaining schedule information, the calendar application then uses the text extraction plugin in the text recognition module to recognize the text in the hotel order screenshot and obtain multiple text information. The intent recognition module can call the word segmentation interface to segment the text information and transmit the segmentation results to the calendar application via a string array.
[0231] The following example illustrates how an electronic device first performs word segmentation processing and then performs schedule information processing.
[0232] Specifically, the execution process of the electronic device is shown in Figure 15:
[0233] 1. The calendar application sends images to the visual classification module.
[0234] After receiving an image, the calendar application in the electronic device sends the image to the visual classification module (or the image classification plugin in the text recognition module).
[0235] 2. The visual classification module identifies the image category.
[0236] The visual classification module identifies the image category, such as whether the image belongs to the first category, the second category, or a specific subcategory within the third category.
[0237] 3. The text recognition module identifies text information in images.
[0238] The calendar application sends images to the text recognition module, enabling the module to identify the text information within the images.
[0239] 4. The entity recognition interface of the intent recognition module filters text information.
[0240] Specifically, the mobile phone can filter text information through the entity recognition interface of the intent recognition module, and then send the filtered text information to the calendar application.
[0241] 5. The calendar application assembles and typesets the filtered text information.
[0242] 6. The intent recognition module segments the text information and returns the segmentation results to the calendar application.
[0243] Specifically, the phone segments the text information using an intent recognition module, stores the segmentation results in an array, and sends them to the calendar application.
[0244] 7. The calendar application sends text information to the schedule information recognition model.
[0245] Specifically, after word segmentation is completed, the calendar application sends text information to the schedule information recognition model.
[0246] 8. The schedule information recognition model analyzes textual information to obtain schedule information.
[0247] 9. The calendar application retrieves word segmentation results and schedule information, and displays them on the user interface.
[0248] Specifically, the calendar application displays the schedule information returned by the schedule information recognition model on the user interface of the screen. The calendar application can also display the schedule title on the user interface. In addition, the calendar application can directly display the word segmentation results on the user interface, or it can display the word segmentation results in response to modifications to the schedule title.
[0249] In some embodiments, after the electronic device generates the schedule title through word segmentation, the schedule title can be edited to modify it and achieve customization. Specifically, in response to the editing operation of the schedule title, the calendar application displays the word segmentation result corresponding to the target image. The calendar application then responds to the selection operation of the target word in the word segmentation result and modifies the schedule title based on the target word.
[0250] The following text, in conjunction with Figures 4 and 16, provides an illustrative explanation of the revised schedule title.
[0251] As shown in Figure 4, after the schedule title 402 corresponding to the ticket order screenshot 101 is displayed on the screen, the electronic device responds to the touch of the edit icon 403 and displays the schedule title modification interface 1601, as shown in Figure 16. This schedule title modification interface 1601 includes all the word segments corresponding to the ticket order screenshot 101, such as "C2201", "Beijing", "Beijing South", "Tianjin", etc. The electronic device then responds to the selection of a word segment, allowing modification of the schedule title in the edit box 1602 within the schedule title modification interface 1601. For example, selecting the word "Beijing South" will automatically change the schedule title from "C2201 train journey" in 402 to "Beijing South, C2201 train journey".
[0252] Other embodiments of this application provide an electronic device that may include a memory and one or more processors. The memory and processors are coupled. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device performs the various functions or steps described in the above method embodiments.
[0253] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.
[0254] This application also provides a computer program product that, when run on a computer, causes the computer to perform various functions or steps performed by the electronic device in the above method embodiments.
[0255] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0256] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0257] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0258] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0259] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0260] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating schedule information, characterized in that, The method, applied to an electronic device, includes: receiving a target image via a calendar application; using the calendar application to call a text recognition module to identify the image category corresponding to the target image and to identify first text information in the target image; using the calendar application to call an intent recognition module to filter the first text information based on the text features corresponding to the image category, obtaining the remaining second text information after filtering; and using the calendar application to send the second text information and the image category to a schedule information recognition model, obtaining the schedule information corresponding to the target image output by the schedule information recognition model.
2. The method according to claim 1, characterized in that, The text features corresponding to the image category include the positional features and / or style features of the text in the images under that image category; The step of calling the intent recognition module through the calendar application to filter the first text information according to the text features corresponding to the image category to obtain the remaining second text information includes: calling the intent recognition module through the calendar application to filter out the text in the first text information that matches the position features and / or style features, and obtaining the second text information.
3. The method according to claim 1 or 2, characterized in that, The text features corresponding to the image category include: keyword information corresponding to the image category; the step of calling the intent recognition module through the calendar application to filter the first text information according to the text features corresponding to the image category to obtain the remaining second text information after filtering includes: calling the intent recognition module through the calendar application to filter out the text in the first text information that matches the keyword information, and obtaining the remaining second text information after filtering.
4. The method according to any one of claims 1-3, characterized in that, The step of calling the text recognition module through the calendar application to identify the image category corresponding to the target image includes: using the text recognition module through the calendar application to perform preliminary identification of the category of the target image to obtain the initial category corresponding to the target image; if the initial category is a first target category, identifying the subcategory corresponding to the target image under the first target category as the image category corresponding to the target image.
5. The method according to any one of claims 1-4, characterized in that, The image categories include any one of the following: chat information images, information notification images, order images, or non-order images.
6. The method according to any one of claims 1-5, characterized in that, The schedule information recognition model includes at least one sub-model; the step of sending the second text information and the image category to the schedule information recognition model through the calendar application to obtain the schedule information corresponding to the target image output by the schedule information recognition model includes: sending the second text information and the image category to the schedule information recognition model through the calendar application, so that the schedule information recognition model determines the target sub-model corresponding to the image category, and analyzes the second text information through the target sub-model to obtain the schedule information corresponding to the target image.
7. The method according to any one of claims 1-6, characterized in that, The method also includes a word segmentation process; The word segmentation process and the process for generating the schedule information are executed in parallel. The steps of the word segmentation process include: calling the word segmentation interface of the intent recognition module through the calendar application to segment the second text information into words, and obtaining the word segmentation result corresponding to the target image.
8. The method according to claim 7, characterized in that, After obtaining the word segmentation result corresponding to the target image, the method further includes: displaying a first interface through the calendar application, the first interface including the schedule information and the schedule title, the schedule title being generated based on the word segmentation result.
9. The method according to claim 8, characterized in that, After the first interface is displayed through the calendar application, the method further includes: the calendar application responding to the editing operation of the event title by displaying the word segmentation result corresponding to the target image; and the calendar application responding to the selection operation of the target word in the word segmentation result by modifying the event title based on the target word.
10. The method according to any one of claims 1-9, characterized in that, After the text recognition module is invoked through the calendar application to identify the image category corresponding to the target image, the method further includes: if the image category is the second target category, sending the target image and the image category to the schedule information recognition model through the calendar application to obtain the schedule information corresponding to the target image; if the image category is not the second target category, performing the step of recognizing the first text information in the target image.
11. An electronic device, characterized in that, The electronic device includes a memory and a processor; the memory and the processor are coupled; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1 to 10.
12. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 10.
13. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 10.