Method for creating schedule information, electronic device and readable storage medium
By identifying pictures text and edges, filtering out interference information, and automatically creating schedule information using the NLP model, it solves the problem of low efficiency of user manual recording schedule information, and realizes efficient and accurate schedule information creation.
Patent Information
- Application Number
- PCT/CN2024/117822
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-09-09
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, users need to manually record the schedule information in third-party social applications in calendar applications, resulting in less efficient creation of schedule information.
The image is recognized through the target recognition model, filters out interference information, extracts schedule information using the NLP model, and automatically creates schedule information.
It improves the efficiency of creating agenda information, reduces user manual operations, and improves the accuracy and user experience of information extraction.
Smart Images

Figure CN2024117822_03072025_PF_FP_ABST
Abstract
Description
Method for creating schedule information, electronic device and readable storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 25, 2023, with application number 202311808698.5 and application name “Method for Creating Schedule Information, Electronic Device and Readable Storage Medium”, and the Chinese patent application filed with the State Intellectual Property Office on December 25, 2023, with application number 202311800855.8 and application name “Method for Creating Schedule Information, Electronic Device and Readable Storage Medium”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of terminal technology, and in particular to a method for creating schedule information, an electronic device, and a readable storage medium. Background Art
[0003] With the rapid development of terminal technology, electronic devices can now install various third-party social applications, such as instant messaging apps. Some messages or service notifications in third-party social applications often include schedule information, such as a conversation in an instant messaging app's chat window that includes a meeting's schedule. In some scenarios, users often need to record schedule information related to third-party social applications on their electronic devices.
[0004] In the related art, users are generally required to manually record schedules in a calendar application, which results in low efficiency in creating schedule information.
[0005] Summary of the Invention
[0006] This application provides a method for creating schedule information, an electronic device, and a readable storage medium, which can solve the problem of low efficiency caused by requiring users to manually create schedule information in related technologies. The technical solution is as follows:
[0007] In a first aspect, a method for creating schedule information is provided, the method comprising:
[0008] In response to a schedule extraction operation on the first image, a first recognition result and a second recognition result are determined through a target recognition model. The first recognition result includes a text recognition result of the first image, and the second recognition result includes an edge recognition result and an image category of the first image. The edge recognition result includes graphic element attribute information or color block attribute information of the first image. According to the image category of the first image, interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first image is filtered out. Based on the filtered first recognition result and the filtered second recognition result, schedule information for the first image is created and displayed. In this way, the user does not need to manually create schedule information in the calendar application, thereby improving the efficiency of creating schedule information.
[0009] As an example of the present application, the target recognition model includes a first optical character recognition (OCR) model, a second OCR model, and multiple edge detection models. The first OCR model can be used to determine the text recognition result of the image, the second OCR model can be used to determine the image category of the image, and different edge detection models can be used to perform edge recognition on images of different image categories. In response to the schedule extraction operation on the first image, the specific implementation of determining the first recognition result and the second recognition result through the target recognition model may include: in response to the schedule extraction operation on the first image, inputting the first image into the first OCR model for recognition processing, and outputting the first recognition result. Inputting the first image into the second OCR model for recognition processing, outputting the image category of the first image, determining the edge detection model corresponding to the image category of the first image from multiple edge detection models, inputting the first image into the determined edge detection model for recognition processing, and outputting the graphic element attribute information or color block attribute information of the first image.
[0010] In this way, after performing text recognition and edge recognition processing on the first picture respectively through the above two branches, the first recognition result and the second recognition result are obtained. Since the text recognition result in the first text recognition result can represent the text content in the first picture, and the edge recognition result in the second recognition result can represent the layout of the first picture, the subsequent schedule information extraction based on the first recognition result and the second recognition result can improve the accuracy of information extraction.
[0011] As an example of this application, picture categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots and other categories of pictures. Instant messaging chat screenshots refer to pictures obtained by taking screenshots of the chat interface in the instant messaging application. Instant messaging notification card screenshots refer to pictures obtained by taking screenshots of the service notification card in the instant messaging application. Order screenshots refer to pictures obtained by taking screenshots of the order interface in the application. Other categories of pictures include pictures other than instant messaging chat screenshots, instant messaging notification card screenshots and order screenshots.
[0012] Among them, when the first picture is an instant messaging chat screenshot, the edge recognition result includes graphic and text element attribute information; when the first picture is an instant messaging notification card screenshot, an order screenshot, or one of other categories of pictures, the edge recognition result includes color block attribute information.
[0013] In this way, by classifying the image categories into the above-mentioned multiple types, in the subsequent schedule information extraction process, the images can be processed in a targeted manner according to the image categories, thereby improving the effectiveness and accuracy of schedule information extraction.
[0014] As an example of the present application, according to the picture category of the first picture, the specific implementation of filtering out interference information that is not related to the schedule in the text recognition results and edge recognition results of the first picture may include: determining the filtering rules corresponding to the picture category of the first picture, different picture categories correspond to different filtering rules, instant messaging chat screenshots correspond to one filtering rule, and instant messaging notification card screenshots and order screenshots correspond to the same filtering rule. According to the filtering rules corresponding to the picture category of the first picture, the interference information that is not related to the schedule in the text recognition results and edge recognition results of the first picture is filtered out. In this way, different filtering rules are used for filtering pictures of different picture categories, and the interference information in pictures of different picture categories can be filtered in a targeted manner, thereby improving the effectiveness of filtering.
[0015] As an example of the present application, the first picture is a screenshot of an instant messaging chat, and the text recognition result includes text line coordinates. Accordingly, according to the filtering rules corresponding to the picture category of the first picture, the specific implementation of filtering out interference information unrelated to the schedule in the text recognition result of the first picture and the edge recognition result can include: filtering out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture. According to the text line coordinates of each text line in the text recognition result of the first picture, filter out the recognition data corresponding to the text lines with a line height less than the target line height in the text recognition result of the first picture and the edge recognition result, and the target line height is the average line height of all text lines in the first picture. Filter out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture.
[0016] In this way, by identifying the data corresponding to the skewed lines, small characters, and chat timestamps in the first picture, interference information unrelated to the schedule can be effectively removed, that is, information that may interfere with the extraction of schedule information can be removed, thereby improving the accuracy of subsequent schedule extraction.
[0017] As an example of the present application, the attribute information of graphic elements in the edge recognition results includes the coordinates and category of the chat elements, and the text recognition results also include text line recognition content. Accordingly, filtering out recognition data corresponding to a chat timestamp in the first image from the text recognition results and edge recognition results of the first image may include: determining chat elements whose category is a chat timestamp from the edge recognition results to obtain at least one first candidate chat element; matching the text line recognition content corresponding to each first candidate chat element from the text recognition results of the first image based on the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition results of the first image; if the target first candidate chat element is determined to be a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determining whether to filter out recognition data corresponding to the target first candidate chat element based on the coordinates of the target first candidate chat element, where the target first candidate chat element is any one of the at least one first candidate chat elements; and if it is determined that the recognition data corresponding to the target first candidate chat element is to be filtered out, filtering out the recognition data corresponding to the target first candidate chat element from the text recognition results and edge recognition results of the first image.
[0018] In this way, the chat element whose category is the chat timestamp is first determined based on the edge recognition result, and then the corresponding text line recognition content is matched from the text recognition result. According to the matched text line recognition content, it is determined through the NLP model whether it is a chat timestamp. Then, it is determined whether it is a chat timestamp based on the position of the chat element. This can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0019] As an example of the present application, a specific implementation of determining whether to filter out recognition data corresponding to a target first candidate chat element based on its coordinates may include: if the target first candidate chat element is determined to be located in the center of a first image based on its coordinates, and the region corresponding to the target first candidate chat element includes a single text line, then determining to filter out recognition data corresponding to the target first candidate chat element. Alternatively, if the target first candidate chat element is determined to be located to the right of the first image based on its coordinates, and the region corresponding to the target first candidate chat element includes a single text line, then determining to filter out recognition data corresponding to the target first candidate chat element if the text line recognition content corresponding to the target first candidate chat element includes only time but not date.
[0020] In this way, by determining whether the first candidate chat element is a chat timestamp based on the position feature and text line feature of the chat timestamp, the accuracy of the determination can be improved.
[0021] As an example of the present application, after filtering out the recognition data corresponding to the chat timestamp in the first image from the text recognition results and edge recognition results of the first image, chat elements unrelated to the chat content are identified from the remaining chat elements in the edge recognition results based on the categories of the chat elements remaining in the edge recognition results, thereby obtaining at least one second candidate chat element. Based on the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition results of the first image, the text line recognition content corresponding to each second candidate chat element is matched from the text recognition results of the first image. The matched text line recognition content and the corresponding text line coordinates are filtered out from the text recognition results of the first image. In this manner, after filtering out the recognition data corresponding to the chat timestamp, small characters, and skewed text lines, the remaining chat elements are further filtered to remove as much interference information as possible from the first image, thereby improving the accuracy of subsequent schedule information extraction.
[0022] As an example of the present application, the first picture is a screenshot of an instant messaging notification card, and the text recognition result includes text line coordinates and text line recognition content. According to the filtering rules corresponding to the picture category of the first picture, the specific implementation of filtering out interference information that is not related to the schedule in the text recognition result of the first picture and the edge recognition result can include: According to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture, the text line recognition content in each color block is matched from the text recognition result of the first picture. For any one of the color blocks, if it is determined that any color block does not include content related to the schedule based on the text line recognition content in any color block, the recognition data corresponding to any color block is filtered out from the edge recognition result and the text recognition result of the first picture. In this way, by filtering out the color blocks that do not include the schedule, so that only the color blocks that include content related to the schedule are subsequently processed, the data processing efficiency can be improved, and the accuracy of schedule information extraction can be improved.
[0023] As an example of the present application, a text recognition result includes text line coordinates and text line recognition content. A specific implementation of creating schedule information for a first image based on the filtered first recognition result and the filtered second recognition result may include: matching the text line recognition content corresponding to each object in the filtered text recognition result of the first image based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first image, where the object is a chat element or a color block. Determining the scene corresponding to the first image based on the text line recognition content corresponding to each object and the image category of the first image. Based on the scene corresponding to the first image, performing text splicing on the text line recognition content corresponding to each object to obtain spliced text. Constructing a prompt based on the spliced text and the scene corresponding to the first image, the prompt includes scene description information that describes the scene corresponding to the first image. Inputting the prompt into a natural language recognition model for processing to extract schedule information from the first image and create the schedule information. In this way, by adding scene description information to the prompt, the NLP model can extract accurate schedule information.
[0024] As an example of the present application, after displaying schedule information, in response to an edit operation on the schedule title of the schedule information, a target interface is displayed, which includes spliced text. In response to a selection operation on the text line content displayed in the target interface, the selected text is entered into the title input box. In response to the editing end operation, the schedule title is modified to the content entered in the title input box. In this way, by displaying the target interface, the user can quickly modify the schedule title by scribbling, thereby improving the user experience.
[0025] In a second aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for creating schedule information as described in the first aspect above is implemented.
[0026] In a third aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is executed on a computer, the computer executes the method for creating schedule information described in the first aspect.
[0027] In a fourth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the method for creating schedule information as described in the first aspect.
[0028] The technical effects obtained by the above-mentioned second, third and fourth aspects are similar to the technical effects obtained by the corresponding technical means in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a schematic diagram showing an application scenario according to an exemplary embodiment;
[0030] FIG2 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0031] FIG3 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0032] FIG4 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0033] FIG5 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0034] FIG6 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0035] FIG7 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0036] FIG8 is a schematic diagram showing a software system of an electronic device according to an exemplary embodiment;
[0037] FIG9 is a schematic diagram showing an implementation framework for creating schedule information according to an exemplary embodiment;
[0038] FIG10 is a flow chart showing a method for creating schedule information according to an exemplary embodiment;
[0039] FIG11 is a schematic diagram showing processing of an instant messaging chat screenshot according to an exemplary embodiment;
[0040] FIG12 is a schematic diagram showing a processing of a screenshot of an instant messaging notification card according to an exemplary embodiment;
[0041] FIG13 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0042] FIG14 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0043] FIG15 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0044] FIG16 is a schematic diagram showing an application scenario according to another exemplary embodiment;
[0045] FIG17 is a schematic diagram showing a process of a method for creating schedule information according to another exemplary embodiment;
[0046] FIG18 is a schematic diagram showing an order screenshot according to an exemplary embodiment;
[0047] FIG19 is a flow chart showing a method for creating schedule information according to another exemplary embodiment;
[0048] Fig. 20 is a schematic diagram showing the architecture of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0050] It should be understood that the “multiple” mentioned in this application refers to two or more. In the description of this application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of this application, words such as “first” and “second” are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as “first” and “second” do not limit the quantity and execution order, and words such as “first” and “second” do not necessarily limit them to be different.
[0051] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0052] Before introducing the method for creating schedule information provided in the embodiment of the present application, a brief explanation of the terms or nouns involved in the embodiment of the present application is first given.
[0053] One-stop office software offers a variety of office functions, including document processing, spreadsheets, presentation creation, project management, calendars, and email. Users can complete multiple office tasks within the same software, improving work efficiency. One-stop office software typically utilizes a unified user interface, making switching between functional modules easier and more consistent.
[0054] KV format: refers to presenting information in a table format (with or without a table). For example, the information presented includes two columns of information, the left side is the topic, and the right side is the specific content corresponding to the topic.
[0055] Global Collection: This means that users can slide down the screen with three fingers to trigger the electronic device to collect information.
[0056] Magic Text: A feature that quickly extracts text from images. You can enable or disable this feature by going to Settings > Smart Assistant > Magic Text and toggling it on or off.
[0057] Prompt: In the embodiment of the present application, it refers to the model input data formed by adding a piece of text or instruction to the input text information when using the machine learning model, that is, it includes user prompt and system prompt, user prompt is the input text information, and system prompt is the added text or instruction. In this way, the machine learning model can be guided to generate more accurate and targeted outputs. The system prompt can be a question, a description, a formatted input, or even some keywords. By reasonably designing the prompt, the machine learning model can be guided to understand the intention of the question, thereby generating a more accurate answer. In short, Prompt is a technology widely used in machine learning models, which can help users solve data bias, improve the controllability of machine learning models and problem modeling capabilities.
[0058] Vertical domain: refers to providing specific services to limited groups, including entertainment, medical care, environmental protection, education, sports and other industries.
[0059] Currently, a variety of third-party social applications can be installed in electronic devices, including, for example, but not limited to, instant messaging (IM) applications, ticket booking applications, etc. For example, IM applications can be WeChat, QQ, DingTalk, Feishu, etc., and ticket booking applications can be Ctrip, Qunar, Tongcheng, Feizhu, etc. Users can chat and book tickets through third-party social applications, or follow the official accounts (or notification accounts) of various industries through third-party social applications to learn about the areas and events they want to pay attention to through the service notification cards published by the official accounts (or notification accounts). In some scenarios, some messages in third-party social applications may involve schedule information, such as one or more messages in the chat window of an instant messaging application involving the relevant schedule of an event. Users usually have the need to record schedule information. In this case, users are generally required to manually create schedule information in a calendar application. However, manually creating schedule information results in cumbersome operations and low efficiency in creating schedule information. To this end, an embodiment of the present application provides a method for creating schedule information, which enables an electronic device to automatically create schedule information, thereby improving the efficiency of creating schedule information.
[0060] As an example of the present application, referring to Table 1, the electronic device may create schedules involved in the following scenarios into schedule information:
[0061] Table 1
[0062] The above-mentioned images and texts include pictures and texts, that is, the method provided in the embodiment of the present application can not only automatically extract schedule information from texts, but also automatically extract schedule information from pictures. Specifically, the scope supported by the method provided in the embodiment of the present application is shown in Table 2:
[0063] Table 2
[0064] According to Table 2, electronic devices can extract schedule information from pictures. Pictures can be screenshots or pictures taken by a camera. The form of screenshots can include but is not limited to full-screen screenshots or area screenshots. The windows involved in the screenshots can be full-screen, split-screen, or floating windows. The source of the screenshots is usually from mobile phones, tablets, and PCs. The form of pictures taken by cameras can be but is not limited to printed text, handwritten text, and artistic text. In addition, electronic devices can also extract schedule information from text. The form of text includes but is not limited to Text and Webview. The format of the text can be plain text, formatted text, or mixed text with pictures. The length of the text can be single paragraph, multiple paragraphs, short text, or long text. The embodiment of this application focuses on the example of input being a picture.
[0065] For ease of understanding, the application scenarios provided by the embodiments of the present application are introduced below.
[0066] In one example, multiple rounds of conversations in a WeChat chat window contain schedule information. When a user wants to create related schedule information, the user can trigger the phone to take a screenshot of the application interface where the chat window is located. For example, the user can double-click the screenshot trigger area on the phone's screen to trigger the screenshot operation, causing the phone to take a screenshot of the interface where the WeChat chat window is located. Subsequently, referring to FIG1 (a), the phone displays a screenshot editing interface U1 and displays the instant messaging chat screenshot obtained after the screenshot is taken, i.e., image p1, in the screenshot editing interface U1. The screenshot editing interface U1 includes a "Share" control. When the user wants to create the schedule information in image p1, they can click the "Share" control. Referring to FIG1 (b), in response to the user's triggering operation on the "Share" control, the phone displays a sharing pop-up window 10. The sharing pop-up window 10 includes a calendar icon 11, and the user can click the calendar icon 11. In response to the user's triggering operation on the calendar icon 11, the mobile phone begins to process the image p1 to extract and create relevant schedule information. For example, see Figure 1 (c). During this process, the mobile phone can display a prompt message "Offline parsing of schedule information" so that the user can know that the schedule is currently being extracted from the image p1. Referring to Figure 1 (d), after the schedule information is successfully created, the mobile phone displays the schedule display interface U2. The schedule display interface U2 includes a schedule display window 12. The schedule display window 12 displays the created schedule information 13. The schedule information 13 includes the schedule title, time, location, etc. In this way, the mobile phone achieves the purpose of automatically creating and displaying schedule information, which can avoid the need for users to manually record and improve the efficiency of creating schedule information.
[0067] As an example of this application, the schedule display window 12 also includes multiple editing controls, allowing users to edit the schedule information 13 created by the phone based on their needs. For example, referring to Figure 2(a), if a user wishes to modify the schedule title, they can click the title editing control 14 in the schedule display window 12. In response to the user clicking the title editing control 14, the phone displays the target interface U3 (also known as the smudge interface) shown in Figure 2(b). Target interface U3 displays information related to the schedule in image p1. The user can modify or fill in the schedule title by smudged on the information displayed on target interface U3. Accordingly, the phone enters the user's smudged content into the title input box 15 of target interface U3. For example, referring to Figure 2(b), if a user wishes to modify the schedule title to "HarmonyOS Open Talk," the user can smudge "HarmonyOS," "Open," and "Talk" in sequence on target interface U3. Accordingly, the phone enters "HarmonyOS," "Open," and "Talk" in sequence into the title input box 15. As shown in Figure 2 (c), after the user finishes smudging, they can trigger the "input" control (or "√" control) on the target interface U3. In response to the user's triggering of the "input" control (or "√" control), the phone resumes displaying the schedule display window 12. The user can now see in the schedule display window 12 that the schedule title has been modified to reflect the content modified by the user's smudging operation, that is, the schedule title has been changed from "KaiTan" to "HarmonyOS KaiTan". In this way, by displaying smudging-enabled schedule-related information on the target interface U3, users can quickly modify the schedule title by smudging, improving the user experience.
[0068] In addition, referring to FIG. 2 (a), the schedule display window 12 also includes a time editing control. When the user wishes to edit the time in the schedule information 13, the user can modify the time based on the time editing control. For example, the user can click on the displayed time to edit. Furthermore, after the user slides down the schedule display window 12, the schedule display window 12 may also provide other editing controls, such as editing controls for the number of repetitions, reminder time, and important reminders. Thus, the user can edit the schedule information based on other editing controls, which is not limited in the present embodiment.
[0069] As an example of the present application, after the user clicks the "√" control in the schedule display window 12, in response to the triggering operation, the mobile phone displays the schedule information 13 in the schedule details area of the calendar application, so that the user can view the schedule information 13 from the schedule details area of the calendar application. As an optional example, after the user clicks the "√" control in the schedule display window 12, the mobile phone can also display the daily schedule information 13 in the form of a card on the desktop, the negative one screen, or the notification center, so that the user can quickly view it later. This embodiment of the present application is not limited to this.
[0070] It should be noted that the above description only uses the example of a user triggering a mobile phone to create schedule information through a sharing portal (i.e., a sharing control). In another example, referring to Figure 1 (a), the mobile phone provides a Magic Text control 00 in the screenshot editing interface U1. When the user needs the mobile phone to create schedule information based on the image p1, they can click the Magic Text control 00, thereby triggering the mobile phone to create and display the schedule information with one click.
[0071] In another example, a user can trigger the phone to create and display schedule information through any portal. For example, referring to Figure 3 (a), after the phone takes a screenshot of the WeChat chat window, the resulting image p1 is automatically saved to the gallery. Thus, when the user wants the phone to automatically create the schedule information for image p1, they can open the screenshot image interface U4 in the gallery, where image p1 is displayed. Referring to Figure 3 (b), the user can trigger the phone to select image p1 and then drag it to the right side of the phone screen. When the user drags it to a certain location, the phone displays applications capable of receiving and processing image p1, such as Calendar, WeChat, and QQ, as shown in Figure 3 (b). The user can then continue dragging image p1 onto the Calendar application and release it. In response to the user's release, the phone begins processing image p1 to extract and create the relevant schedule information. Referring to Figure 3 (c), after the phone creates the schedule information, it displays it in the schedule display window 12.
[0072] In another example, the user can circle only the content related to the schedule in picture p1, and then trigger the mobile phone to extract schedule information for the circled part. For example, referring to Figure 4 (a), the user can circle part of the content in picture p1. For example, a control for triggering the circle operation can be provided in the screenshot editing interface. After the user triggers the control, the user can circle the picture p1. In response to the user's circle operation, the mobile phone selects the part circled by the user. As an example, the mobile phone can take a screenshot of the area circled by the user, as shown in Figure 4 (b). Afterwards, the user can trigger the mobile phone to extract schedule information for the selected area through interactive entrances such as the "Share" control, and create and display schedule information related to the content in the area. In this way, by supporting users to circle picture p1, the data processing capacity of the mobile phone can be reduced, thereby improving the efficiency of creating schedule information.
[0073] It should be noted that the interactive portals used by the user in the above application scenarios to trigger the mobile phone to extract schedule information from image p1 are merely exemplary. In some embodiments, the mobile phone can also be triggered to extract schedule information from image p1 through other interactive portals, such as through interactive portals such as global favorites, which are not limited in this embodiment of the present application.
[0074] The above application scenarios are merely exemplary. Furthermore, the mobile phone can also extract schedule information from certain types of images in other scenarios. For example, referring to FIG5 , a forwarded image p2 (which can be a screenshot or a photo) is present in a chat window, and schedule information is contained in image p2. To extract the schedule information from image p2, the user can drag image p2 to the right side of the phone's screen, as shown in FIG5 (b). When the user drags the image to a certain location, the phone displays applications capable of receiving and processing image p2, such as Calendar, WeChat, and QQ, as shown in FIG5 (b). The user drags image p2 onto the calendar application and releases it. In response to the user's release, the phone automatically processes image p2 to extract and create relevant schedule information. Referring to FIG5 (c), after the phone creates the schedule information, it displays it in the schedule display window 12.
[0075] Refer to Figure (c) in Figure 5. When the mobile phone creates schedule information multiple times, multiple schedule information can be displayed in the schedule display window 12. When the schedule display window 12 does not display all the schedule information, the user can trigger the mobile phone to display the hidden schedule information by sliding the schedule display window 12 left and right.
[0076] Furthermore, the phone supports not only dragging images from a chat window to any portal, but also text. For example, if a chat window contains chat text and schedule information, and the user needs to create the schedule information from the chat text, they can drag the chat text to the calendar application as shown in Figure 5. The phone will then extract the schedule information from the chat text, create it, and display it.
[0077] In another example, the mobile phone can also extract schedule information from a screenshot of an instant messaging notification card. For example, referring to Figure 6 (a), shown in the figure is a schematic diagram of a screenshot of an instant messaging notification card (i.e., picture p3) according to an exemplary embodiment. Picture p3 is obtained by taking a screenshot of a notification card issued by a service notification in WeChat, which includes schedule information. When the user wants the mobile phone to create and display the schedule information in picture p3, the mobile phone can be triggered to extract the schedule information according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information for picture p3 through the sharing entrance. Accordingly, the mobile phone extracts the schedule information based on picture p3, and then creates or displays the relevant schedule information. For example, referring to Figure 6 (b), the mobile phone creates and displays schedule information 60. Schedule information 60 includes subject, ticket collection number, departure time and date, location, etc.
[0078] In another example, the mobile phone can also extract schedule information from an order screenshot, where the order screenshot can be a screenshot of a hotel order, a train ticket order, an airplane ticket order, or the like. For example, referring to FIG7 (a), the figure shows a schematic diagram of an order screenshot (i.e., image p4) according to an exemplary embodiment. Image p4 is a screenshot of the interface where the train ticket order is located, and image p4 includes schedule information. When the user wants the mobile phone to create and display the schedule information in image p4, the mobile phone can be triggered according to the operation process described above. For example, the user can trigger the mobile phone to extract schedule information from image p4 through the Magic Text portal. Accordingly, the mobile phone extracts schedule information based on image p4 and then creates or displays relevant schedule information. For example, the displayed schedule information is shown as 70 in FIG7 (b). Schedule information 70 includes the subject, train number, departure time and arrival time, date, location, etc.
[0079] It should be noted that the above application scenarios are exemplary and do not limit the application scenarios of the method provided in the embodiments of this application. In another embodiment, the mobile phone can also extract schedule information from other types of pictures, including pictures other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots. Of course, other types of pictures can be pictures obtained through screenshots or pictures taken with a camera.
[0080] It should also be noted that the above description is based on an example of a mobile phone as an electronic device. The electronic devices involved in the embodiments of the present application may also be sports cameras (GoPro), digital cameras, tablet computers, desktop computers, laptop computers, handheld computers, notebook computers, vehicle-mounted equipment, ultra-mobile personal computers (UMPCs), netbooks, etc., and the embodiments of the present application are not limited to this.
[0081] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to exemplify the software system of the electronic device.
[0082] Figure 8 is a block diagram of a software system for an electronic device provided in an embodiment of the present application. Referring to Figure 8 , the layered architecture divides the software into several layers, each with clear roles and divisions of labor. Layers communicate with each other via software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime layer, the system layer, and the kernel layer, from top to bottom.
[0083] The application layer may include a series of application packages, see Figure 8, the application package may include a calendar and other applications, for example, other applications may include instant messaging, ticket booking, camera, gallery, call, map, navigation, Bluetooth, music, video, short message and other applications.
[0084] In addition, as an example of the present application, the application layer also includes a schedule management service and a model management service. The schedule management service can be used to provide services for the calendar, or to be called by the calendar. For example, the schedule management service can create schedule information for the calendar and store data such as schedule information. The schedule management service may include a schedule creation service (which may be called: intelligent parsing and processing service) and a schedule database (such as: Calednar Provider schedule database). The schedule management service can create schedule information through the schedule creation service and store the created schedule information through the schedule database. The model management service (which may be called: MagicLive large model service) can be used to provide various models for the schedule management service to call when needed.
[0085] In one example, the model management service can provide a natural language processing (NLP) model, an object recognition model, and a personal behavior feature model. The NLP model can be used to recognize prompts to determine schedule information. The NLP model can be run using natural language units (NLU). In some examples, the NLP model can not only recognize prompts but also extract keywords from a text, such as time and location. The object recognition model can be used to perform text recognition and edge recognition on images, and can also be used to determine the image category. In one example, the object recognition model includes a first optical character recognition (OCR) model, a second OCR model, and an edge detection model. The first OCR model can be used for text recognition, the second OCR model can be used to determine the image category, and the edge detection model can be used for edge recognition on images. There can be multiple edge detection models, and different edge detection models can be used for edge recognition of images of different categories. The personal behavior feature model can be used to determine a user profile based on historical user behavior data.
[0086] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. As shown in Figure 8, the application framework layer may include a window manager, content provider, view system, telephony manager, resource manager, notification manager, etc.
[0087] The Android Runtime consists of core libraries and a virtual machine (VM). The Android runtime is responsible for scheduling and management of the Android system. The core library consists of two parts: one for Java-based functions and the other for the Android core library. The application layer and application framework layer run in the VM. The VM executes Java files from the application layer and application framework layer as binary files. The VM is responsible for performing functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0088] The system library can include multiple functional modules, such as: surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0089] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0090] The electronic device can implement the method for creating schedule information provided by the embodiment of the present application through the interaction of the above-mentioned multiple modules. For example, refer to Figure 9, which is a schematic diagram of a method implementation framework shown according to an exemplary embodiment. In implementation, the electronic device can pass the image to be processed to the calendar application through interactive entrances such as sharing, any door, global collection or Magic text. Among them, the picture can be a screenshot of a ticket or air ticket order, or a screenshot of a hotel, catering, or entertainment order, or a screenshot of a sports or health order, or a screenshot of a WeChat applet, web, etc., or a screenshot of a service notification card of an IM notification number, or a screenshot of an IM chat conversation, etc. As an example and not a limitation, screenshots of train and air ticket orders can come from applications such as 12306, China Railway or Ctrip; screenshots of hotel, catering and entertainment orders can come from applications such as Ctrip, Qunar, Tongcheng, Meituan, Fliggy, Dianping, Damai, etc.; screenshots of sports and health orders can come from applications such as Keep, Lianduoduo, and registration platforms; screenshots of WeChat applets, web, etc. can include performances, shopping sprees, marathons, etc.; screenshots of service notification cards of IM notification numbers can include notification cards issued by third parties such as hospitals, scenic spot tickets, educational institutions, and insurance companies through IM notification numbers; screenshots of IM chat conversations can include work arrangements, invitations, educational tasks, and other content.
[0091] After receiving the image, the calendar application requests the target recognition model to perform text recognition and edge recognition on the image. Text recognition can extract text information from the image, while edge recognition can identify chat elements or color blocks within the image. For example, referring to Figure 9, image p1 is passed to the target recognition model through the interactive portal. After the target recognition model performs text recognition, the resulting text recognition result is shown as 90 in Figure 9. Furthermore, edge recognition by the target recognition model can identify the coordinates and categories of each chat element in image p1.
[0092] Afterwards, the electronic device can filter the text recognition results and edge recognition results based on the filtering rules to remove interference information in the text recognition results and edge recognition results that is not related to the schedule. As an example of the application, the preset filtering rules may include but are not limited to at least one of the following rules: 1. Discarding dense text blocks. 2. Discarding small characters. 3. Discarding crooked lines. 4. Discarding floating text on the attached figure. 5. Discarding the back in the upper left corner of the picture and the "+" symbol in the picture. 6. Discarding color blocks that are not related to the schedule. For example, after filtering the text recognition results based on the filtering rules, the content shown in 91 in Figure 9 can be removed. In addition, after filtering the edge recognition results based on the filtering rules, the coordinates and categories of chat elements that are not related to the schedule can be removed.
[0093] If the text recognition results are not filtered but directly input into the NLP model for schedule information extraction, in this case, as shown in Table 3, the recognition success rate of the NLP model is low, usually only reaching about 30%. The reason is that there are problems such as interference information, format line breaks, and layout information loss. Therefore, the embodiment of the present application filters the text recognition results before using the NLP model to extract schedule information, which can improve the accuracy of the subsequent NLP model in extracting schedule information.
[0094] Table 3
[0095] After filtering, the electronic device constructs a prompt (i.e., prompt) based on the text recognition results and edge recognition results remaining after filtering, and then inputs the prompt into the NLP model to extract the schedule information through the NLP model. The NLP model outputs the schedule information based on the input prompt. For example, 92 in Figure 9 shows part of the content processed by the NLP model. After that, by post-processing the schedule field of the schedule information, the schedule information to be displayed can be obtained. For example, the processed schedule information is shown as 93 in Figure 9. After that, the electronic device can display the schedule information.
[0096] In one example, the post-processing of the schedule field may include, but is not limited to, at least one of the following: 1. Processing according to the reminder time rule. 2. Processing according to the tomorrow start time rule. 3. Processing according to the details rule. 4. Processing based on the title merge pre-fill rule.
[0097] The reminder time rule includes setting the reminder time of the schedule information to be earlier than the preset time of the time extracted by the NLP model. The preset time can be set according to needs. For example, if the preset time is 30 minutes, if the time extracted by the NLP model is 8:30, then the reminder time of the schedule information can be set to 8:30. In addition, the reminder time rule also includes repeated reminders. For example, if the schedule information extracted by the NLP model is for grabbing numbers on Monday, Tuesday and Wednesday, the electronic device will set the reminder time on Monday, Tuesday and Wednesday for the schedule information, instead of setting only one reminder time.
[0098] Tomorrow's start time rule means that if the time extracted by the NLP model spans days, months, or years, the time after the span will be supplemented. For example, if the date extracted by the NLP model is December 5, the start time is 23:00, and the end time is 00:30, the electronic device can add the end time as December 6 00:30 in the schedule information.
[0099] The detail rule refers to adjusting the layout and font size of the created schedule information according to the size of the screen of the electronic device so that it can be correctly and clearly displayed in the schedule details area of the calendar application.
[0100] The title merging and pre-population rules include selecting the longest title from multiple titles corresponding to the same time as the schedule information title. Furthermore, the title merging and pre-population rules also include using the designated schedule title corresponding to the image scene as the schedule information title. For example, if the image scene is about making an appointment to get a number for a medical consultation, if the schedule title extracted by the NLP model is "Get a Number" or "Get a Number," the schedule title can be standardized to "Register for a Medical Consultation." The designated schedule titles corresponding to different scenarios can be pre-set as needed.
[0101] It should be noted that the above-mentioned post-processing of the schedule field is only exemplary. In another example, the post-processing of the schedule field may also include but is not limited to at least one of time similarity processing, discarding empty results, risk control, and cleaning non-natural language titles. Time similarity processing means that if the time in the output schedule information is earlier than the current system time of the electronic device, the time closest to the time in the schedule information is determined based on the current system time, and the determined time is determined as the time in the schedule information. For example, if the time in the schedule information is Tuesday, and the current system time is Wednesday, it is recorded in the schedule information as Tuesday of the next week. Discarding empty results means discarding the returned empty fields. Risk control refers to the management and control of sensitive words. Cleaning non-natural language titles means that if the schedule title does not include verbs, a verb can be added to the schedule title or the schedule title can be set by default according to the scene corresponding to the picture.
[0102] Next, the method for creating schedule information provided by the embodiment of the present application is described in detail with reference to FIG10. Referring to FIG10, the method may include the following steps:
[0103] S1001: The calendar application receives a picture L to be processed.
[0104] The image L may be an image in bitmap format.
[0105] The calendar application receives the image L delivered by the interactive portal. As mentioned above, the interactive portal can be a sharing portal, an Anydoor portal, etc. For example, referring to FIG1 (a), when the user submits the image L to the calendar application through the sharing control, the interactive portal is the sharing portal.
[0106] S1002: The calendar application sends a schedule creation instruction to the schedule management service, and the schedule creation instruction carries a picture L.
[0107] The schedule creation instruction is used to instruct to create and display related schedule information based on the image L.
[0108] S1003: The schedule management service sends the picture L to the first OCR model in the model management service.
[0109] In implementation, the schedule management service calls the first OCR model in the model management service and sends the image L to the first OCR model so that the first OCR model performs text recognition processing.
[0110] S1004: The first OCR model determines a first recognition result of the image L.
[0111] As an example of the present application, the first recognition result includes a text recognition result of the image L, and the text recognition result includes text block coordinates, text line coordinates, and text line recognition content.
[0112] As an example, the first recognition result also includes target indication information, which can be used to indicate whether the image L input into the first OCR model is a screenshot or a photographed image. Exemplarily, the target indication information can be a first identifier, a second identifier, or a third identifier. The first identifier is used to indicate that the image L input into the first OCR model is a screenshot, the second identifier is used to indicate that the image L input into the first OCR model is a photographed image and is a photograph of a document, and the third identifier is used to indicate that the image L input into the first OCR model is other photographed images, such as images of advertisements, road signs, magazines, etc. The first identifier, the second identifier, and the third identifier can be set as needed, for example, the first identifier is F1, the second identifier is F2, and the third identifier is F3.
[0113] That is, after receiving the image L, the first OCR model recognizes the image L and outputs a first recognition result of the image L.
[0114] S1005: The first OCR model sends a first recognition result to the schedule management service.
[0115] As an example, after receiving the first recognition result, the schedule management service may cache the first recognition result.
[0116] S1006: The schedule management service sends the picture L to the second OCR model in the model management service.
[0117] In one example, after the schedule management service receives a schedule creation instruction sent from the calendar application, in addition to sending the image L to the first OCR model for text recognition processing, it can also call the second OCR model in the model management service and send the image L to the second OCR model to determine the image category through the second OCR model, that is, the operation of S1006 and the operation of S1003 can be executed in parallel.
[0118] Because the method provided in the embodiments of this application can extract schedule information from images of different image categories, and images of different image categories have different content layouts, electronic devices process images of different image categories differently. Therefore, in implementation, after receiving an image L to be processed, the schedule management service not only performs text recognition using the first OCR model, but also inputs image L into the second OCR model to determine the image category of image L.
[0119] As an example of this application, the image categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other categories of images. Instant messaging chat screenshots refer to images obtained by taking screenshots of the chat interface in the instant messaging application; instant messaging notification card screenshots refer to images obtained by taking screenshots of the service notification card in the instant messaging application; order screenshots refer to images obtained by taking screenshots of the order interface in the application, such as screenshots of train ticket orders, air ticket orders, and hotel orders; other categories of images include images other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots. Other categories of images can be screenshots or photographs.
[0120] S1007: The second OCR model determines the image category of the image L.
[0121] After receiving the image L, the second OCR model recognizes the image L and outputs the image category of the image L.
[0122] S1008: The second OCR model sends the image category of image L to the schedule management service.
[0123] S1009: The schedule management service sends picture L to the edge detection model corresponding to the picture category of picture L.
[0124] As an example of the present application, a plurality of edge detection models are provided in the model management service, and the plurality of edge detection models can be pre-trained, and different edge detection models can perform edge recognition on images of different image categories. In one example, the plurality of edge detection models include a first edge detection model and a second edge detection model. The first edge detection model can be used to perform edge recognition on instant messaging chat screenshots to determine the coordinates and categories of chat elements in the instant messaging chat screenshots; the second edge detection model can be used to perform edge recognition on images other than instant messaging chat screenshots, for example, the second edge detection model can be used to perform edge recognition on instant messaging notification card screenshots, order screenshots, or other categories of images to determine the coordinates of color blocks of other images, and further determine the categories of color blocks.
[0125] Different edge detection models can be obtained by iteratively training an initial training model based on image training samples of corresponding image categories. The image training samples can be pre-obtained through edge annotation as needed. The initial training model can be set as needed. As an example and not limitation, the initial training model can be a recurrent neural network (RNN), etc.
[0126] After receiving the image category of image L, the schedule management service determines an edge detection model corresponding to the image category of image L from multiple edge detection models. For example, if image L is p1 in Figure 1(a), i.e., an instant messaging chat screenshot, the edge detection model corresponding to the image category of image L is determined to be the first edge detection model from the multiple edge detection models. If image L is p3 in Figure 6(a) (i.e., an instant messaging notification card screenshot) or p4 in Figure 7(a) (i.e., an order screenshot), the edge detection model corresponding to the image category of image L is determined to be the second edge detection model from the multiple edge detection models. The schedule management service then invokes the determined edge detection model and sends image L to the edge detection model to request edge recognition of image L.
[0127] S1010: The edge detection model performs edge recognition processing on the image L and outputs the edge recognition result.
[0128] The edge recognition result includes the attribute information of graphic elements or color blocks.
[0129] In one example, if the image category of image L is instant messaging chat screenshot, the edge detection model corresponding to the image category of image L is a first edge detection model. After image L is input into the first edge detection model for processing, the output edge recognition result includes attribute information of graphic elements. The graphic element attribute information includes attribute information of chat elements in the instant messaging chat screenshot. The attribute information of chat elements may include the coordinates and category of the chat elements. Exemplarily, chat elements include avatars, titles, nicknames, chat content, chat timestamps, usernames, and designated identifiers, where designated identifiers include "+" identifiers. For example, referring to FIG11 , after image L is input into the first edge detection model, the first edge detection model may determine that the chat elements in image L include the multiple items indicated by the dashed boxes in FIG11 . Each chat element has corresponding coordinates and categories. For example, the coordinates of a chat element are the coordinates of the four corners of the area where the chat element is located, and the category is avatar. Optionally, the attribute information of each chat element may also include a sequence number of the chat element. The sequence number of each chat element in image L may be set by default based on the order of the chat elements in image L.
[0130] In another example, when picture L is not an instant messaging chat screenshot, for example, a screenshot of an instant messaging notification card, an order screenshot, or another category of picture, the edge detection model corresponding to the picture category of picture L is the second edge detection model. After picture L is input into the second edge detection model, the second edge detection model performs color block segmentation, and the output edge recognition result includes color block attribute information. Exemplarily, the color block attribute information includes the coordinates and color block number of the color block. For example, referring to Figure 12, when picture L is a screenshot of an instant messaging notification card, after picture L is input into the second edge detection model, the second edge detection model can determine that the color blocks in picture L include the multiple items marked by the dotted box in Figure 12, and each color block corresponds to its own coordinates, for example, the coordinates of the four corners of the color block.
[0131] S1011: The edge detection model sends edge recognition results to the schedule management service.
[0132] As an example, after receiving the edge recognition result, the schedule management service may cache the edge recognition result.
[0133] It's worth noting that after performing text recognition and edge recognition on image L through the two branches described above, a first recognition result and a second recognition result are obtained. The first recognition result includes the text recognition result and target indication information, and the second recognition result includes the edge recognition result and image category. Because the text recognition result can represent the text content in image L, and the edge recognition result can represent the layout of image L, subsequent schedule information extraction based on the first and second recognition results can improve the accuracy of information extraction. The specific implementation can be seen in the following steps.
[0134] S1012: When the image category of the image L is an instant messaging chat screenshot, the schedule management service filters the text recognition result and the edge recognition result according to the first filtering rule.
[0135] Different image categories correspond to different filtering rules. In implementation, the schedule management service determines the corresponding filtering rules based on the image category of image L, and then filters the text recognition results and edge recognition results of image L according to the determined filtering rules to filter out interference information that is not related to the schedule.
[0136] In one example, an instant messaging chat screenshot corresponds to a first filtering rule, which may include filtering out identification data corresponding to skewed lines, small text, and chat timestamps. A skewed line refers to a text line corresponding to a chat element with an inclination angle greater than a preset angle, which can be set as needed, such as 10 degrees. Small text refers to a text line corresponding to a chat element with a height less than a target height, which can be the average height of all text lines in image L. The chat timestamp is a timestamp indicating the chat time. Optionally, the first filtering rule may also include filtering out dense text blocks, floating text on the image, back in the upper left corner, and a "+" in the lower right corner. For example, referring to FIG. 11 , FIG. 11 identifies chat elements in image L that need to be filtered, including chat timestamp 1101, small text 1102, skewed lines 1103, and "+" 1004, according to an exemplary embodiment.
[0137] As an example, in the implementation of filtering out the recognition data corresponding to the skewed lines, the inclination angle of each text line can be calculated based on the text line coordinates in the text recognition results, so as to determine which text lines are skewed lines, and then delete the recognition data corresponding to the skewed lines, such as deleting the text line coordinates and text line recognition content of the skewed lines.
[0138] As an example, in the implementation of filtering out the recognition data corresponding to small characters, the line height of each text line in the image L can be determined based on the text line coordinates in the text recognition results, thereby filtering out the recognition data corresponding to the text lines whose line height is less than the target line height, such as filtering out the text line coordinates and text line recognition content of the text lines whose line height is less than the target line height.
[0139] As an example, in an implementation of filtering out recognition data corresponding to a chat timestamp, chat elements categorized as chat timestamps can be determined based on edge recognition results to obtain at least one first candidate chat element. Based on the coordinates of each of the at least one first candidate chat elements and the text line coordinates in the text recognition results of image L, the text line recognition content corresponding to each first candidate chat element is matched from the text recognition results of image L. If the target first candidate chat element is determined to be a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, a determination is made based on the coordinates of the target first candidate chat element whether to filter out recognition data corresponding to the target first candidate chat element, where the target first candidate chat element is any one of the at least one first candidate chat elements. If the determination is to filter out recognition data corresponding to the target first candidate chat element, the recognition data corresponding to the target first candidate chat element is filtered out from the text recognition results and edge recognition results of image L.
[0140] As previously mentioned, the edge recognition results of instant messaging chat screenshots include chat element categories. Therefore, based on the chat element categories in the edge recognition results, chat elements with a chat timestamp category can be filtered out to obtain at least one first candidate chat element. Because the edge recognition results do not include text line recognition content, meaning the text content corresponding to each chat element cannot be determined, in some cases, the edge detection model may mistakenly classify chat elements that are not chat timestamps as chat timestamps. Therefore, to minimize the filtering of chat elements that are not chat timestamps, after determining at least one first candidate chat element, the text line recognition content corresponding to each first candidate chat element can be matched from the text recognition results of image L based on the coordinates of each first candidate chat element and the text line coordinates of image L. For example, for any first candidate chat element, based on the coordinates of the first candidate chat element and the text line coordinates of image L, the text line recognition content of at least one text line located within the area corresponding to the first candidate chat element can be determined from the text recognition results of image L, thereby matching the text line recognition content corresponding to the first candidate chat element. The schedule management service can then use the NLP model to determine whether the matched text line identification content is a timestamp. For example, the schedule management service can send the matched text line identification content to the NLP model through getEntity() to request the NLP model to identify whether the text line identification content is a timestamp. If the NLP model determines that the text line identification content is a timestamp, it can be further determined that the first candidate chat element is likely a chat timestamp. Otherwise, if the NLP model determines that the text line identification content is not a timestamp, it can be determined that the first candidate chat element is not a chat timestamp.
[0141] Since the position of the chat timestamp in an instant messaging chat screenshot is generally fixed, for example, it is usually displayed in the center of the WeChat chat window. Therefore, when the NLP model is used to perform text line content recognition to determine that a first candidate chat element may be a chat timestamp, it is possible to further determine whether this first candidate chat element is a chat timestamp based on the coordinates of this first candidate chat element, thereby improving the accuracy of chat timestamp recognition.
[0142] As an example of the present application, for any one of the at least one first candidate chat element, determining whether to filter out the identification data corresponding to the first candidate chat element based on the coordinates of the first candidate chat element may include the following two specific implementations:
[0143] The first case: when it is determined based on the coordinates of the first candidate chat element that the first candidate chat element is located in the middle of the image L, and the area corresponding to the first candidate chat element includes a single text line, it is determined to filter out the recognition data corresponding to the first candidate chat element.
[0144] In most instant messaging apps, chat timestamps are typically displayed in the center of the chat window, and the chat timestamp only includes a single line of text line recognition content, namely, the timestamp. Therefore, if it is determined that the first candidate chat element is located in the center of image L, and the area corresponding to the first candidate chat element only includes a single text line, and since the NLP model has determined that the text line recognition content is a timestamp, the first candidate chat element can be determined to be a chat timestamp, and the recognition data corresponding to the first candidate chat element can be filtered out.
[0145] The second case: when it is determined that the first candidate chat element is located on the right side of the image L according to the coordinates of the first candidate chat element, and the area corresponding to the first candidate chat element includes a single text line, if the text line recognition content corresponding to the first candidate chat element only includes time but not date, it is determined to filter out the recognition data corresponding to the first candidate chat element.
[0146] Because in the chat windows of some instant messaging applications (such as the chat interface forwarded in WeChat), the chat timestamp may also be displayed on the right, and the chat timestamp only includes a line of text line recognition content. In addition, the chat timestamp only includes time but not date. Therefore, if it is determined that this first candidate chat element is located on the right side of the image L, and the area corresponding to this first candidate chat element only includes a single text line, it can be determined whether the text line recognition content corresponding to this first candidate chat element includes a date. For example, the NLP model can be used to determine whether the text line recognition content includes a date. If it is determined that the date is not included, it can be determined that this first candidate chat element is a chat timestamp, that is, it can be determined that the recognition data corresponding to this first candidate chat element can be filtered out. Of course, if it is determined that the date is included, it can be determined not to be filtered out.
[0147] In one example, for the second case, it is not necessary to determine whether the text line identification content corresponding to the first candidate chat element only includes time but not date. As long as it is determined that the first candidate chat element is located on the right side of the image L and the area corresponding to the first candidate chat element includes a single text line, the schedule management service can determine to filter out the identification data corresponding to the first candidate chat element.
[0148] If the above process determines that a first candidate chat element is a chat timestamp, the schedule management service deletes the identification data corresponding to this first candidate chat element from the text recognition results and edge recognition results of image L. For example, the schedule management service deletes the text line coordinates and text line recognition content corresponding to this first candidate chat element from the text recognition results of image L, and deletes the coordinates and category corresponding to this first candidate chat element from the edge recognition results of image L. Of course, if the above process determines that a first candidate chat element is not a chat timestamp, the schedule management service does not filter out the identification data corresponding to this first candidate chat element.
[0149] It is worth mentioning that the chat element category is first determined to be a chat timestamp based on the edge recognition results, and then the corresponding text line recognition content is matched from the text recognition results. Based on the matched text line recognition content, the NLP model is used to determine whether it is a chat timestamp. After that, it is determined whether it is a chat timestamp based on the position of the chat element. This can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0150] It should be noted that the first filtering rule described above is merely exemplary. When instant messaging chat screenshots originate from different instant messaging applications, the layout of their chat elements is typically different, and the chat elements may also be different. This may result in differences in the interference information in instant messaging chat screenshots from different instant messaging applications. For example, these may typically include the several possible scenarios shown in Table 4. Therefore, in another example, the first filtering rule may further include other rules for filtering out content unrelated to the chat content.
[0151] Table 4
[0152] To effectively filter out interference information from instant messaging chat screenshots from different instant messaging applications, a first filtering rule can be set based on the union of possible interference information shown in Table 4, thereby ensuring that interference information can be effectively removed regardless of the type of instant messaging chat screenshot being processed. Exemplarily, the first filtering rule can also include filtering out text, usernames, and designated identifiers in avatars. In implementation, after filtering out the recognition data corresponding to the chat timestamp in image L from the text recognition results and edge recognition results of image L, the schedule management service can determine chat elements unrelated to the chat content from the remaining chat elements in the edge recognition results based on the categories of the chat elements remaining in the edge recognition results, thereby obtaining at least one second candidate chat element. Based on the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition results of image L, the text line recognition content corresponding to each second candidate chat element is matched from the text recognition results of image L. The matched text line recognition content and the corresponding text line coordinates are deleted from the text recognition results of image L.
[0153] Exemplarily, in the implementation of filtering out text in the avatar, the schedule management service can determine the chat element whose category is the avatar based on the edge recognition result, and then match the text line recognition content corresponding to the chat element in the current recognition result of the image L based on the coordinates of the chat element. If there is a matching text line recognition content, the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content are deleted from the text line recognition result of the image L, thereby deleting the text in the avatar.
[0154] Exemplarily, in the implementation of filtering out user names, the schedule management service can determine the chat element whose category is the user name based on the edge recognition result, and then match the text line recognition content corresponding to the chat element in the text recognition result of image L based on the coordinates of the chat element. If there is a matching text line recognition content, the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content are deleted from the text line recognition result of image L, thereby deleting the user name.
[0155] For example, in an implementation that filters out a specified identifier, the schedule management service can determine, based on the edge recognition results, chat elements whose categories are the specified identifiers. The schedule management service can then match the text line identification content corresponding to the chat element from the edge recognition results of image L. If the text line identification content is the specified identifier, such as "+," the identification data corresponding to the chat element can be deleted from the text line recognition results of image L, such as the text line coordinates and text line identification content corresponding to the chat element. The schedule management service can also delete the identification data corresponding to the chat element from the edge recognition results, such as the coordinates and category of the chat element.
[0156] It should be noted that the above description is based on the example that picture L is a screenshot of an instant messaging chat. In another example, if picture L is not a screenshot of an instant messaging chat, such as a screenshot of an instant messaging notification card or an order screenshot, the schedule management service filters the interference information based on the second filtering rule. In one example, in the implementation of filtering based on the second filtering rule, the schedule management service can match the text line recognition content in each color block from the text recognition result of picture L according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of picture L. For any one of the color blocks, if it is determined based on the text line recognition content in any one of the color blocks that any one of the color blocks does not include content related to the schedule, such as time and place, then the recognition data corresponding to any one of the color blocks is filtered out from the edge recognition result and the text recognition result of picture L. For specific implementation, please refer to the embodiment shown in Figure 17.
[0157] In one example, an NLP model can be used to determine whether a color block includes time and place. For example, the text line recognition content in the color block can be sent to the NLP model one by one to request the NLP model to determine whether the time and place are included.
[0158] If image L is a screenshot of an instant messaging notification card, Table 4 shows that its possible interference information includes small text, skewed lines, and a central timestamp. Therefore, in one example, before filtering based on the second filtering rule, the text line recognition content of image L and the identification data corresponding to the skewed lines, small text, and central timestamp in the edge recognition results can be filtered out, and then filtered out based on the second filtering rule. When filtering the central timestamp, the corresponding text line recognition content can be matched, and the text line recognition content and the coordinates of the color block can be used to determine whether it is a central timestamp.
[0159] As an example of this application, before filtering, the schedule management service can also determine whether there are dense text blocks based on the text block coordinates in the text recognition results of image L. If dense text blocks exist, a prompt message can be displayed in the calendar application to inform the user of the dense text blocks so that the user can re-capture image L as needed. If there are no dense text blocks, the schedule management service performs the filtering operation.
[0160] Or in another example, during the filtering process, if the schedule management service determines that there are dense text blocks based on the text block coordinates in the text recognition results of the image L, the text line coordinates and text line recognition content in these text blocks are deleted.
[0161] It's worth noting that without filtering, the subsequent NLP model can be unable to accurately extract schedule information. For example, for image p1 in Figure 9, the time might be extracted as the chat timestamp "20:37," and the address might be extracted as "Ningyuan Road, No. 5, Danling Street, Haidian District, Beijing." In this embodiment of the present application, after determining the first and second recognition results, filtering out interference information in image L can improve the accuracy of subsequent schedule information extraction.
[0162] S1013: The schedule management service matches the text line recognition content of the remaining chat elements from the filtered text recognition results based on the filtered edge recognition results.
[0163] The schedule management service matches the text line recognition content corresponding to each chat element in the filtered text recognition result according to the coordinates of each chat element in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.
[0164] When picture L is not an instant messaging chat screenshot, the schedule management service matches the text line recognition content in each filtered color block from the filtered text recognition result based on the coordinates of each color block in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.
[0165] S1014: The schedule management service identifies the content of the remaining text lines of each chat element and the image category of the image L, and determines the scene corresponding to the image L.
[0166] The scene corresponding to the picture L refers to the scene involved in the text content in the picture L.
[0167] When it is determined that picture L is an instant messaging chat screenshot based on its picture category, if it is further determined to be a chat conversation based on the filtered text line recognition content of each chat element, then the scene corresponding to picture L can be determined to be an IM chat scene.
[0168] In addition, when the image category of picture L is other images, such as an instant messaging notification card screenshot or an order screenshot, the schedule management service can also combine the text line recognition content of each matched color block to determine the scene corresponding to picture L, such as a service notification card scene or a high-speed rail travel order scene.
[0169] S1015: The schedule management service splices the chat dialogues based on the filtered text line recognition content of each chat element based on the scene corresponding to the image L.
[0170] During implementation, the schedule management service performs line breaks, carriage returns, and splicing on the filtered text lines of each chat element based on the scene corresponding to image L. Because chat content is often quite casual, for example, a complete sentence may be sent as multiple messages. If the scene corresponding to image L is determined to be an IM chat, the schedule management service can splice these multiple messages into a single sentence. Therefore, splicing chat conversations based on the scene corresponding to image L can make the spliced text more similar to natural language text. As an example and not a limitation, the text after chat conversation splicing is consistent with the content displayed in the smudge interface.
[0171] In one example, the spliced chat content includes the conversation type, title, nickname, conversation content, etc. The nickname can be customized. For example, the chat content can be spliced in the following format:
[0172] Conversation topic: IM chat
[0173] Title: xx group chat
[0174] Nickname A: xxx
[0175] Nickname B: xxx
[0176] Nickname A: xxxx
[0177] .......
[0178] It should be noted that S1014 to S1016 are optional operations. In another example, the schedule management service may also perform chat dialogue splicing on the filtered text line recognition content of each chat element according to the image category of the image L.
[0179] In addition, when picture L is another picture, such as a screenshot of an instant messaging notification card or an order screenshot, the schedule management service will splice the text line recognition content in each filtered color block according to the scene corresponding to picture L. Its implementation can be seen in the splicing of chat conversations.
[0180] S1016: The schedule management service constructs a prompt based on the scene corresponding to the image L and the spliced text.
[0181] Different image categories correspond to different prompt construction templates. When constructing a prompt, the schedule management service can use the prompt construction template corresponding to the image category of image L. As an example of this application, the constructed prompt includes scene description information, which is used to indicate the scene corresponding to image L. That is, the schedule management service adds the scene description information of image L when constructing the prompt.
[0182] It's worth noting that if you don't include contextual descriptions when constructing prompts, the NLP model will likely misinterpret each time and the verbs associated with it as a single calendar item during subsequent recognition. This can lead to multiple calendar items being extracted, resulting in inaccurate calendar information. Therefore, to ensure accurate calendar information recognition, the calendar management service adds contextual descriptions to the constructed prompts, allowing the NLP model to accurately extract a single calendar item.
[0183] For example, taking the case where picture L is a screenshot of an instant messaging chat, the prompt constructed by the schedule management service may be:
[0184] <|Human|> The following content may be an IM chat conversation. IM chat conversations contain the fields <title, time, location, participants>. Each field in the event is output as a line in JSON format, and no additional response is given. \n\nKaiTan
[0185] HarmonyOS Architecture Evolution and Key Technologies
[0186] HDC Together
[0187] Time: 09:00-16:30 October 23
[0188] Location: Microsoft Asia Pacific R&D Group Building
[0189] No. 5 Danling Street, Haidian District, Beijing
[0190] <|Moss|>
[0191] S1017: The schedule management service sends a prompt to the NLP model.
[0192] After the schedule management service constructs the prompt, it calls the NLP model and sends the constructed prompt to the NLP model for recognition to extract the schedule information. For example, the schedule management service can request the NLP model to extract the schedule information through extractinformation().
[0193] As mentioned above, the NLP model not only has the ability to recognize schedule information, but also has the ability to extract keywords. Since there may be a certain delay in the NLP model's recognition of prompts, if the recognition time is long, it will affect the user experience. Therefore, in some examples, the schedule management service can also send the spliced text to the NLP model line by line at the behavioral granularity to request the NLP model to extract keywords from the spliced text. In this way, when the recognition time is long, the extracted keywords can be used to construct schedule information. For example, the schedule management service can instruct the NLP model to extract keywords through getEntity(). For example, the NLP model can be specified to extract time and place keywords, that is, the module can be specified as time and location.
[0194] It should be noted that the embodiments of this application are described using the example of deploying an NLP model in an electronic device. In another example, the NLP model can also be deployed in the cloud, which can provide an interface for the electronic device to call the NLP model. In this way, when the NLP model is needed, the schedule management service can call the NLP model through the provided interface. This embodiment of this application is not limited to this.
[0195] S1018: The NLP model determines the schedule information based on the prompt.
[0196] In one example, the schedule information output by the NLP model includes: {"data":"[{'Title':'KaiTan','Time','Start Time':'09:00','End Time':'16:30','Date':'October 23','Location':'Microsoft Asia Pacific R&D Group Building, No. 5 Danling Street, Haidian District, Beijing'}]"}.
[0197] In one example, when the schedule management service further sends the concatenated text to the NLP model to instruct the NLP model to extract keywords, the schedule management service may also extract corresponding keywords.
[0198] S1019: The NLP model sends the schedule information to the schedule management service.
[0199] For example, the NLP model may send schedule information to the schedule management service in JSON format.
[0200] Furthermore, if the schedule management service extracts keywords, the extracted keywords are also sent to the schedule management service.
[0201] S1020: The schedule management service performs schedule field post-processing on the schedule information.
[0202] For the post-processing of the schedule field of the schedule information, please refer to the above.
[0203] In one example, if the NLP model returns schedule information within a specified time period, the schedule management service performs schedule field post-processing based on the schedule information fed back by the NLP model to create the schedule information. Since the schedule information is identified based on prompts, the accuracy of schedule information creation can be improved. If the NLP model does not return schedule information within the specified time period, it means that the schedule information recognition has timed out. In this case, the schedule management service can construct the schedule information based on the keywords extracted by the NLP model to minimize the problem of schedule information creation being stuck. The specified time period is set according to needs, for example, the specified time period can be 5 seconds.
[0204] It should be noted that when creating schedule information based on keywords, schedule fields can also be post-processed and then created according to a preset template, for example, including content such as subject, time, and location.
[0205] In one example, before creating a schedule, the schedule management service can also call a personal behavior model to query the user's historical behavior data, such as historical locations. The personal behavior model then returns the historical behavior data. The schedule management service can then predict the user's likely destinations based on this historical behavior data. It then combines this with the schedule information provided by the NLP model to create the final schedule, for example, by adding the predicted address information to the schedule.
[0206] S1021: The schedule management service displays the processed schedule information in the calendar application.
[0207] For example, when the picture L is a screenshot of an instant messaging chat, the schedule information displayed by the electronic device is as shown in 13 in FIG. 1 (d).
[0208] In one example, before displaying the schedule information, a confirmation request notification may be displayed first. After receiving a confirmation display instruction triggered by the user based on the confirmation request notification, the schedule management service displays the schedule information in the calendar application.
[0209] As an example of the present application, the electronic device also supports users to edit the displayed schedule information. For example, it supports users to modify the schedule title of the schedule information, for example, see the embodiment shown in Figure 2. In this process, when the user triggers the electronic device to display the target interface, it is necessary to display the smudgeable text in the target interface. For this purpose, the schedule management service can also request the NLP model to perform word combination processing on the spliced text, such as combining the two words "I" and "we" into "we", so that the text can be displayed in the target interface according to the combined vocabulary, which is convenient for users to smudge. In implementation, the schedule management service can call the NLP model through getWordSegment() and request the NLP model to perform word combination processing. In addition, the schedule management service can also call the NLP model through getWordSegment() and request the NLP model to determine the subject of the spliced text, such as specifying the NLP model to extract the subject entity (such as meeting, dinner). In this way, the schedule management service can display the subject in the target interface.
[0210] In an embodiment of the present application, after receiving a schedule extraction operation for picture L, a first recognition result and a second recognition result are determined by a target recognition model in response to the schedule extraction operation. The first recognition result includes the text recognition result of picture L, and the second recognition result includes the edge recognition result and the picture category of picture L. The edge recognition result includes the graphic element attribute information or color block attribute information of picture L. According to the picture category of picture L, interference information unrelated to the schedule in the text recognition result and the edge recognition result is filtered out, and then based on the filtered first recognition result and the second recognition result, the schedule information of picture L is created and displayed. In this way, the user does not need to manually enter the schedule information item by item in the calendar application, thereby improving the efficiency of creating schedule information.
[0211] The above embodiment is mainly explained by taking the example of picture L being an instant messaging chat screenshot. As mentioned above, picture L may also be other pictures, for example, it may be an order screenshot, for example, an order screenshot may be a screenshot of a hotel, train ticket, airplane ticket, etc. Next, the method for creating schedule information provided by the embodiment of the present application is introduced by taking the example of picture L being an order screenshot. First, the relevant application scenarios are introduced:
[0212] For example, taking a screenshot of a train ticket order, when a user wants to create related schedule information, the user can trigger the phone to take a screenshot of the application interface containing the train ticket order. For example, when the phone displays the application interface containing the train ticket order, the user can double-click the screen in the screenshot trigger area to trigger the screenshot operation, causing the phone to take a screenshot of the application interface containing the train ticket order. Subsequently, referring to FIG13 (a), the phone displays a screenshot editing interface U11, and displays the screenshot of the train ticket order obtained after the screenshot is taken, i.e., image p5, in the screenshot editing interface U11. The screenshot editing interface U11 includes a "Share" control. When the user wants to create the schedule information in image p5, they can click the "Share" control. Referring to FIG13 (b), in response to the user triggering the "Share" control, the phone displays a sharing pop-up window 130. The sharing pop-up window 130 includes a calendar icon 131, and the user can click the calendar icon 131. In response to the user's triggering operation on the calendar icon 131, the mobile phone begins to process the image p5 to extract and create relevant schedule information. For example, see Figure 13 (c). During this process, the mobile phone can display a prompt message "Offline parsing of schedule information" so that the user can know that the schedule is currently being extracted from the image p5. See Figure 13 (d). After the schedule information is successfully created, the mobile phone displays the schedule display interface U12. The schedule display interface U12 includes a schedule display window 132. The schedule display window 132 displays the created schedule information 133. The schedule information 133 includes the schedule title, train number, departure time and arrival time, origin and destination, etc. In this way, the mobile phone achieves the purpose of automatically creating and displaying schedule information, which can avoid the need for users to manually record and improve the efficiency of creating schedule information.
[0213] As an example of the present application, the schedule display window 132 also includes multiple editing controls, and the user can also edit the schedule information 133 created by the mobile phone based on the multiple editing controls as needed. For example, referring to Figure 14 (a), when the user wants to modify the schedule title, he can click the title editing control 134 in the schedule display window 132. In response to the user's click operation on the title editing control 134, the mobile phone displays the target interface U13 (which can be called the smear interface) as shown in Figure 14 (b), and the target interface U13 displays information related to the schedule in picture p5. In this way, the user can modify or fill in the schedule title by smearing on the information displayed on the target interface U13. Accordingly, the mobile phone enters the user's smeared content in the title input box 135 of the target interface U13. For example, referring to Figure 14 (b), when the user wants to change the schedule title to "Xiao Bo's Order Details", the user can sequentially fill in "Xiao Bo," "Order," and "Details" in the target interface U13. Accordingly, the mobile phone enters "Xiao Bo," "Order," and "Details" in the title input box 135. Referring to Figure 14 (c), after the user finishes filling in the content, the "Input" control (or "√" control) of the target interface U13 can be triggered. In response to the user's triggering operation on the "Input" control (or "√" control), referring to Figure 14 (d), the mobile phone resumes displaying the schedule display window 132. At this time, the user can see from the schedule display window 132 that the schedule title has been modified to the content modified by the user through the filling operation, that is, the schedule title has been changed from "Ticket Purchase Successful Notification" to "Xiao Bo's Order Details." In this way, by displaying information related to the schedule that can be filled in the target interface U13, the user can quickly modify the schedule title by filling in the content, thereby improving the user experience.
[0214] In addition, referring to FIG. 14 (a), the schedule display window 132 also includes a time editing control. When the user wishes to edit the time in the schedule information 133, the user can modify the time based on the time editing control. For example, the user can click on the displayed time to edit. Furthermore, after the user slides down the schedule display window 132, the schedule display window 132 may also provide other editing controls, such as editing controls for the number of repetitions, reminder time, and important reminders. Thus, the user can edit the schedule information based on other editing controls, which is not limited in the present embodiment.
[0215] As an example of the present application, after the user clicks the "√" control in the schedule display window 132, in response to the triggering operation, the mobile phone displays the schedule information 133 in the schedule details area of the calendar application, so that the user can view the schedule information 133 from the schedule details area of the calendar application. As an optional example, after the user clicks the "√" control in the schedule display window 132, the mobile phone can also display the daily schedule information 133 in the form of a card on the desktop, the negative one screen, or the notification center, so that the user can quickly view it later. This embodiment of the present application is not limited to this.
[0216] It should be noted that the above description is based on the example of a user triggering the phone to create schedule information through a sharing portal (i.e., a sharing control). In another example, see Figure 13 (a), the phone provides a Magic Text control 1300 in the screenshot editing interface U11. When the user needs the phone to create schedule information based on image p5, they can click the Magic Text control 1300 to trigger the phone to create and display the schedule information with one click.
[0217] In another example, a user can trigger the phone to create and display schedule information through any entrance. For example, referring to Figure 15 (a), after the phone takes a screenshot of the application interface containing the order screenshot, the resulting image p5 is automatically saved to the gallery. Thus, when the user wants the phone to automatically create the schedule information for image p5, they can open the screenshot image interface U14 in the gallery, where image p5 is displayed. Referring to Figure 15 (b), the user can trigger the phone to select image p5 and then drag it to the right side of the phone screen. When the user drags it to a certain location, the phone displays applications capable of receiving and processing image p5, such as Calendar, WeChat, and QQ, as shown in Figure 15 (b). The user can then continue dragging image p1 onto the Calendar application and release it. In response to the user's release, the phone begins processing image p5 to extract and create the relevant schedule information. Referring to FIG. 15 ( c ), after the mobile phone creates the schedule information, the schedule information is displayed in the schedule display window 132 .
[0218] In another example, the user can circle only the content related to the schedule in image p5, and then trigger the mobile phone to extract schedule information for the circled part. For example, referring to Figure 16 (a), the user can circle part of the content in image p5. For example, a control for triggering the circle operation can be provided in the screenshot editing interface U11, and after triggering the control, the user can circle the image p5. In response to the user's circle operation, the mobile phone selects the part circled by the user. As an example, the mobile phone can take a screenshot of the area circled by the user, as shown in Figure 16 (b). Afterwards, the user can trigger the mobile phone to extract schedule information for the selected area through interactive entrances such as the "Share" control, and create and display schedule information related to the content in the area. In this way, by supporting the user to circle image p5, the data processing capacity of the mobile phone can be reduced, thereby improving the efficiency of creating schedule information.
[0219] It should be noted that the interactive portals used by the user in the above-mentioned application scenarios to trigger the mobile phone to extract schedule information from image p5 are merely exemplary. In some embodiments, the mobile phone can also be triggered to extract schedule information from image p5 through other interactive portals, such as through interactive portals such as global favorites, which are not limited in this embodiment of the present application.
[0220] Next, the method flow of creating schedule information by an electronic device in the application scenarios shown in Figures 13 to 16 is introduced. Referring to Figure 17, the method may include some or all of the following contents:
[0221] For S1701 to S1711 , reference may be made to S1001 to S1011 in the embodiment shown in FIG. 10 .
[0222] S1712: When the image category of the image L is an order screenshot, the schedule management service filters the recognition data corresponding to the skewed lines, small characters, and the middle timestamp from the text recognition results and the edge recognition results of the image L.
[0223] Different image categories correspond to different filtering rules. In implementation, the schedule management service determines the corresponding filtering rules based on the image category of image L, and then filters the text recognition results and edge recognition results of image L according to the determined filtering rules to filter out interference information that is not related to the schedule.
[0224] As an example of the present application, when picture L is a screenshot of an order, the schedule management service can first filter the recognition data corresponding to the skewed lines, small characters, and the middle timestamp from the text recognition results and edge recognition results of picture L. Among them, the skewed lines refer to text lines with an inclination angle greater than a preset angle, and the preset angle can be set according to needs, for example, the preset angle is 10 degrees or 15 degrees; small characters (usually carried in the attached drawings) refer to text lines with a line height less than the target line height, and the target line height can refer to the average line height of all text lines in picture L. Small characters are generally less than 10dp; the middle timestamp is a timestamp used to indicate the time when the order was generated, usually located in the middle of picture L.
[0225] As an example, in the implementation of filtering out the recognition data corresponding to the skewed lines, the inclination angle of each text line can be calculated based on the text line coordinates in the text recognition results, so as to determine which text lines are skewed lines, and then delete the recognition data corresponding to the skewed lines, such as deleting the text line coordinates and text line recognition content of the skewed lines.
[0226] As an example, in the implementation of filtering out the recognition data corresponding to small characters, the line height of each text line in the image L can be determined based on the text line coordinates in the text recognition results, and the recognition data corresponding to the text lines with a line height less than the target line height can be filtered out, thereby filtering out the recognition data corresponding to small characters, such as filtering out the text line coordinates and text line recognition content of the text lines with a line height less than the target line height.
[0227] As an example, in the implementation of filtering out the recognition data corresponding to the middle timestamp, the color blocks whose category is a timestamp can be determined from the edge recognition results based on the categories of each color block in the edge recognition results to obtain at least one candidate color block. Based on the coordinates of each candidate color block in the at least one candidate color block and the text line coordinates in the text recognition results of the image L, the text line recognition content corresponding to each candidate color block is matched from the text recognition results of the image L. For any candidate color block, if the text line recognition content corresponding to the any candidate color block is a timestamp, the position of the any candidate color block in the image L is determined based on the coordinates of the any candidate color block. If the recognition data corresponding to the any candidate color block is determined to be filtered out based on the position of the any candidate color block, the recognition data corresponding to the any candidate color block is filtered out from the text recognition results and the edge recognition results of the image L.
[0228] Specifically, since the edge recognition results of the order screenshot include the categories of color blocks, color blocks with the timestamp category can be filtered out based on the categories of the color blocks in the edge recognition results to obtain at least one candidate color block. Since the edge recognition results do not include text line recognition content, that is, the text content corresponding to each color block cannot be known, in some possible cases, the edge detection model may misjudge the category of color blocks that are not timestamps. Therefore, in order to avoid filtering out color blocks that are not timestamps as much as possible, after determining at least one candidate color block, the text line recognition content corresponding to each candidate color block can be matched from the text recognition results of image L based on the coordinates of each candidate color block and the text line coordinates of image L. For example, for any candidate color block, based on the coordinates of this candidate color block and the text line coordinates of image L, the text line recognition content of at least one text line located within this candidate color block is determined from the text recognition results of image L, thereby matching the text line recognition content corresponding to this candidate color block. Afterwards, the schedule management service can use the NLP model to determine whether the matched text line recognition content is a timestamp. For example, the schedule management service can send the matched text line recognition content to the NLP model through getEntity() to request the NLP model to identify whether the text line recognition content is a timestamp. If the NLP model determines that the text line recognition content is a timestamp, it can be further determined that this candidate color block may be a mid-time stamp. Otherwise, if the NLP model determines that the text line recognition content is not a timestamp, it can be determined that this candidate color block is not a mid-time stamp.
[0229] Furthermore, since the position of the timestamp in the order screenshot is generally fixed, for example, it is located in the middle of the picture L, that is, most of them are middle timestamps, and the middle timestamp only includes one line of text line recognition content, that is, it only includes a timestamp. Therefore, when the text line content recognition is performed through the NLP model to determine that the text line recognition content in a candidate color block is a timestamp, if it is determined based on the coordinates of this candidate color block that this candidate color block is located in the middle of the picture L, and the area corresponding to this candidate color block includes a single text line, then this candidate color block is determined to be the middle timestamp, that is, its corresponding recognition data is determined to be filtered out.
[0230] When the above process determines that a candidate color block is a middle timestamp, the schedule management service deletes the identification data corresponding to this candidate color block from the text recognition results and edge recognition results of image L. For example, the text line coordinates and text line recognition content corresponding to this candidate color block are deleted from the text recognition results of image L, and the coordinates and category of this candidate color block are deleted from the edge recognition results of image L.
[0231] It should be noted that the embodiment of the present application is explained by taking the case where the timestamp in the order screenshot is the middle timestamp as an example. In some embodiments, the timestamp in the order screenshot may also be located on the right or left side, that is, not the middle timestamp. In this case, for any candidate color block, after determining that the text line recognition content in the any candidate color block is a timestamp, if it is determined based on the coordinates of the any candidate color block that the any candidate color block is located at the right position (or left position) of the image L, and the any candidate color block includes a single text line, then it is determined that the recognition data corresponding to the any candidate color block is filtered out.
[0232] It is worth mentioning that the color block category is first determined to be a timestamp based on the edge recognition results, and then the corresponding text line recognition content is matched from the text recognition results. Based on the matched text line recognition content, the NLP model is used to determine whether it is a timestamp. Then, based on the position of the color block, it is determined whether it is a middle timestamp (or right timestamp, or left timestamp). This can improve the accuracy of timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0233] It should be noted that the above filtering rules are only exemplary. In the case where the order screenshots come from different applications, the information layout is usually different, and the information included may also be different, so that the interference information in the order screenshots of different applications may be different. For example, referring to Table 5, the interference information that may exist in the hotel order from the Meituan application is the small words and crooked lines in the bottom picture, while the interference information that may exist in the hotel order from the Qunar application is the small words or crooked lines in the lower right corner of the picture. In order to effectively filter out the interference information in the order screenshots from different applications, in addition to filtering the small words, crooked lines, and middle timestamps, other interference information that is not related to the schedule can also be filtered, so as to ensure that no matter which order screenshot is processed, the interference information can be effectively removed. Exemplarily, the filtering rules corresponding to the order screenshots can also include filtering out specified identifiers, keyboard text, etc., which is not limited in the embodiments of the present application.
[0234] Table 5
[0235] It should be noted that S1712 is an optional operation.
[0236] S1713: The schedule management service filters out recognition data corresponding to color blocks that do not contain schedule-related content in the text recognition results and edge recognition results of the image L.
[0237] In one example, the schedule management service can match the text line recognition content in each color block from the filtered text recognition results based on the coordinates of each color block in the filtered edge recognition results and the text line coordinates in the filtered text recognition results. For any one of the color blocks, if it is determined based on the text line recognition content in the color block that the color block does not include schedule-related content, such as time and location, the recognition data corresponding to the color block is filtered out from the filtered edge recognition results and the filtered text recognition results, such as filtering out the color block attribute information of the color block from the filtered edge recognition results, and filtering out the text line coordinates and text line recognition content included in the color block from the filtered text recognition results.
[0238] In one example, the schedule management service can use the NLP model to determine whether the color block includes time and place. For example, the text line recognition content in the color block can be sent to the NLP model one by one to request the NLP model to determine whether the time and place are included.
[0239] As an example of the present application, before the schedule management service executes the S1713 operation, it can also determine whether the number of color blocks included in the image L exceeds the specified number based on the color block attribute information in the edge recognition result. If the number of color blocks included in the image L does not exceed the specified number, it means that there are not a large number of color blocks in the image L. In this case, the electronic device is usually able to perform recognition processing on the image L. At this time, the schedule management service executes the S1714 operation. Otherwise, if the number of color blocks included in the image L exceeds the specified number, it means that the image L includes a large number of color blocks. In this case, the schedule information may not be accurately extracted, so filtering processing may not be performed. For example, a prompt message can be displayed in the calendar application to prompt the user that there are too many color blocks, thereby guiding the user to re-capture the image L as needed. Among them, the specified number can be set according to needs. For example, the specified number can be 10, and this embodiment of the present application is not limited to this.
[0240] As an example of the present application, before filtering, the schedule management service can also determine whether there are dense text blocks based on the text block coordinates in the text recognition results of image L. If dense text blocks exist, for example, the number of text blocks exceeds a preset threshold, a prompt message can be displayed in the calendar application to inform the user of the dense text blocks so that the user can re-capture image L as needed. If no dense text blocks exist, the schedule management service performs the filtering operation.
[0241] Or in another example, during the filtering process, if the schedule management service determines that there are dense text blocks based on the text block coordinates in the text recognition results of the image L, the text line coordinates and text line recognition content in these text blocks are deleted.
[0242] It is worth mentioning that if filtering is not performed, the subsequent NLP model may be unable to accurately extract schedule information. For example, for the image in Figure 18, the start time may be extracted as "2023.12.02", and the schedule subject may be taken as "ordering food". In the embodiment of the present application, after determining the first and second recognition results, filtering out interference information in image L can improve the accuracy of subsequent schedule information extraction. In addition, removing color blocks unrelated to the schedule can improve the efficiency of schedule information extraction.
[0243] It should also be noted that the above description uses the example of image L being an order screenshot. In another example, if image L is a screenshot of an instant messaging notification card, similar filtering can also be performed. In another example, if image L is a screenshot of an instant messaging chat, the schedule management service filters out recognition data corresponding to interfering information such as small characters, skewed lines, and chat timestamps from the edge recognition results and text recognition results of image L. For specific implementation, see S1712.
[0244] S1714: The schedule management service determines the scene corresponding to the picture L based on the text line recognition content of the remaining color blocks and the picture category of the picture L.
[0245] The scene corresponding to the picture L refers to the scene involved in the text content in the picture L, such as a ticket booking scene, an IM chat scene, a hotel booking scene, a medical appointment scene, etc.
[0246] As an example, when it is determined that picture L is an order screenshot based on its image category, the order scenario can be further determined based on the text line recognition content of each filtered color block, such as whether it is a ticket booking scenario or a hotel booking scenario. Ticket booking scenarios include ordering train tickets or plane tickets.
[0247] In addition, when the image category of picture L is other images, such as an instant messaging notification card screenshot or an instant messaging chat screenshot, the schedule management service can also determine the scene corresponding to picture L based on the matched text line recognition content, such as determining the IM chat scene.
[0248] S1715: When the scene corresponding to the image L is a ticket booking scene, for any one of the remaining color blocks, if the any one of the remaining color blocks includes multiple text lines, the management service will perform text splicing by line in the future.
[0249] During implementation, the schedule management service performs line breaks, carriage returns, spaces, and concatenation on the text lines identified in each filtered color block based on the scene corresponding to image L. If the scene corresponding to image L is a ticket booking scenario, that is, image L is a screenshot of a train ticket order or a plane ticket order, as shown in Table 5, the start time, end time, origin, and destination in a train ticket order screenshot or a plane ticket order screenshot are typically concatenated by column. For example, a train ticket output from the 12306 ticket booking app would be "Beijing South 06:00 C2551 Binhai 06:56," while a train ticket output from the national railway ticket booking app would be "C2551 06:00 Beijing South, full journey 56 minutes, 06:56 Binhai." However, this does not conform to natural language text and can easily lead to inaccurate recognition by subsequent NLP models. To this end, the schedule management service splices text by row. For example, the train ticket information from the 12306 ticket booking application is converted from columns to rows and spliced into "Beijing South C2551 Binhai 06:00 06:56", and the train ticket information from the railway ticket booking application is converted from columns to rows and spliced into "C2551 06:00 06:56 Beijing South full journey 56 minutes Binhai", making the spliced text closer to natural language text.
[0250] S1716: During the splicing process, if any of the remaining color blocks contains interactive button text, the schedule management service deletes the interactive button text.
[0251] The interactive button text is the text line identification content in the interactive button. The interactive button corresponds to a single text line, the text line length is less than the length threshold, and the text line identification content is a verb. For example, the third color block from the top to the bottom in Figure 12, "Order Meal," "Constitute Risk," and "QR Code Ticket Check," are interactive button text. During the splicing process, these interactive button texts are deleted, that is, they are not spliced, thereby avoiding schedule information extraction errors to a certain extent. The length threshold can be set as needed.
[0252] Of course, it should be noted that the embodiment of the present application is explained by taking the filtering of interactive button text during the splicing process as an example. In another example, it can also be filtered out at other times, for example, it can be performed after filtering color blocks that are not related to the schedule, or it can be performed before filtering color blocks that are not related to the schedule. The embodiment of the present application does not limit this.
[0253] In this way, after the text lines in each color block are identified and spliced, the remaining color blocks can be spliced into a complete spliced text according to the order in which they are arranged in the image L. For example, referring to FIG18 , the spliced text is shown as 1801 in FIG18 . As an example and not a limitation, the spliced text is consistent with the content displayed in the smear interface.
[0254] It should be noted that S1714 to S1716 are optional operations. In another example, the schedule management service may also perform text splicing on the text line recognition contents of each filtered color block according to the image category of the image L.
[0255] It should also be noted that the above description uses the scenario corresponding to image L as a ticket booking scenario. In another example, the scenario corresponding to image L may also be other scenarios, such as hotel booking or IM chat. Accordingly, the schedule management service can similarly perform text splicing based on the scenario using pre-set splicing rules. For example, in an IM chat scenario, a user may send a single sentence in multiple messages. After determining the scenario of image L, the schedule management service can splice the multiple messages sent by the user into a single sentence based on the IM chat scenario.
[0256] S1717: The schedule management service constructs a prompt based on the scene corresponding to the image L and the spliced text.
[0257] Different image categories correspond to different prompt construction templates. When constructing a prompt, the schedule management service can use the prompt construction template corresponding to the image category of image L. As an example of this application, the constructed prompt includes scene description information, which is used to indicate the scene corresponding to image L. That is, the schedule management service adds the scene description information of image L when constructing the prompt.
[0258] It's worth noting that if you don't include contextual descriptions when constructing prompts, the NLP model will likely misinterpret each time and the verbs associated with it as a single calendar item during subsequent recognition. This can lead to multiple calendar items being extracted, resulting in inaccurate calendar information. Therefore, to ensure accurate calendar information recognition, the calendar management service adds contextual descriptions to the constructed prompts, allowing the NLP model to accurately extract a single calendar item.
[0259] In one example, a prompt can be constructed based on the layout order of each remaining color block in image L, the scene corresponding to image L, and the concatenated text. In another example, a prompt can be constructed for each remaining color block based on the concatenated text of each remaining color block and the scene corresponding to image L, i.e., one prompt is constructed for each color block.
[0260] For example, taking picture L as an order screenshot, the prompt constructed by the schedule management service can be:
[0261] <|Human|> The following content may be a travel order, containing the fields <'Title', 'Start Time', 'End Time', 'Recurrence', 'Location', 'URL', 'Convener', and 'Participants'>. Each field in the event is output as a single line in JSON format, and no additional response is provided. \n\nBeijing South Railway Station 07:21 December 5, 2023 Wuwei Railway Station 11:41 December 5, 2023 <|Moss|>.
[0262] S1718: The schedule management service sends a prompt to the NLP model.
[0263] After the schedule management service constructs the prompt, it calls the NLP model and sends the constructed prompt to the NLP model for recognition to extract the schedule information. For example, the schedule management service can request the NLP model to extract the schedule information through extractinformation().
[0264] As mentioned above, the NLP model not only has the ability to recognize schedule information, but also has the ability to extract keywords. Since there may be a certain delay in the NLP model's recognition of prompts, if the recognition time is long, it will affect the user experience. Therefore, in some examples, the schedule management service can also use color blocks as the granularity, and send the spliced text of each remaining color block to the NLP model one by one to request the NLP model to extract keywords from the spliced text of each remaining color block. In this way, when the recognition time is long, the extracted keywords can be used to construct schedule information. As an example, the schedule management service can instruct the NLP model to extract keywords through getEntity(). For example, the NLP model can be specified to extract keywords such as time and place, that is, the module can be specified as time and location.
[0265] It should be noted that the embodiments of this application are described using the example of deploying an NLP model in an electronic device. In another example, the NLP model can also be deployed in the cloud, which can provide an interface for the electronic device to call the NLP model. In this way, when the NLP model is needed, the schedule management service can call the NLP model through the provided interface. This embodiment of this application is not limited to this.
[0266] S1719: The NLP model determines the schedule information based on the prompt.
[0267] In an example, the schedule information output by the NLP model is: {"data":"[{'Title':'Go to Wuwei','Start Time':'December 5, 2023 07:21','End Time':'December 5, 2023 11:41','Location':'Nanwuwei, Beijing'}]"}.
[0268] In one example, when the schedule management service sends the concatenated text of each color block to the NLP model, the schedule management service can also extract keywords from each color block.
[0269] S1720: The NLP model sends schedule information to the schedule management service.
[0270] For example, the NLP model may send schedule information to the schedule management service in JSON format.
[0271] Furthermore, if the schedule management service extracts keywords, the extracted keywords in each color block are also sent to the schedule management service.
[0272] S1721: The schedule management service performs schedule field post-processing on the schedule information.
[0273] For the post-processing of the schedule field of the schedule information, please refer to the above.
[0274] In one example, if the NLP model returns schedule information within a specified time period, the schedule management service performs schedule field post-processing based on the schedule information fed back by the NLP model to create the schedule information. Since the schedule information is identified based on prompts, the accuracy of schedule information creation can be improved. If the NLP model does not return schedule information within the specified time period, it means that the schedule information recognition has timed out. In this case, the schedule management service can construct the schedule information based on the keywords extracted by the NLP model to minimize the problem of schedule information creation being stuck. The specified time period is set according to needs, for example, the specified time period can be 5 seconds.
[0275] It should be noted that when creating schedule information based on keywords, schedule fields can also be post-processed and then created according to preset templates, such as including schedule title, train number, departure and arrival time, origin and destination, etc.
[0276] In one example, before creating a schedule, the schedule management service can also call a personal behavior model to query the user's historical behavior data, such as historical locations. The personal behavior model then returns the number of historical behaviors. The schedule management service can then predict the user's likely destinations based on this historical behavior data. It then combines this with the schedule information provided by the NLP model to create the final schedule, for example, by adding the predicted address information to the schedule.
[0277] S1722: The schedule management service displays the processed schedule information in the calendar application.
[0278] Exemplarily, when picture L is a screenshot of an order, the schedule information displayed on the electronic device is as shown in 133 in (d) of FIG. 13 .
[0279] In one example, before displaying the schedule information, a confirmation request notification may be displayed first. After receiving a confirmation display instruction triggered by the user based on the confirmation request notification, the schedule management service displays the schedule information in the calendar application.
[0280] As an example of the present application, the electronic device also supports users to edit the displayed schedule information. For example, it supports users to modify the schedule title of the schedule information, for example, see the embodiment shown in Figure 14. In this process, when the user triggers the electronic device to display the target interface, it is necessary to display the smudgeable text in the target interface. For this purpose, the schedule management service can also request the NLP model to perform word combination processing on the spliced text, such as combining the two words "I" and "we" into "we", so that the text can be displayed in the target interface according to the combined vocabulary, which is convenient for users to smudge. In implementation, the schedule management service can call the NLP model through getWordSegment() and request the NLP model to perform word combination processing. In addition, the schedule management service can also call the NLP model through getWordSegment() and request the NLP model to determine the subject of the spliced text, such as specifying the NLP model to extract the subject entity (such as meeting, dinner). In this way, the schedule management service can display the subject in the target interface.
[0281] In an embodiment of the present application, after receiving a schedule extraction operation for picture L, a first recognition result and a second recognition result are determined by a target recognition model in response to the schedule extraction operation. The first recognition result includes the text recognition result of picture L, and the second recognition result includes the edge recognition result and the picture category of picture L. The picture category of picture L is an order screenshot, and the edge recognition result includes the color block attribute information of picture L. According to the picture category of picture L, interference information unrelated to the schedule in the text recognition result and the edge recognition result is filtered out, and then based on the filtered first recognition result and the second recognition result, the schedule information of picture L is created and displayed. In this way, the user does not need to manually create the schedule information in the order screenshot in the calendar application, which improves the efficiency of creating schedule information.
[0282] Next, the implementation process of the method for creating schedule information provided by the embodiment of the present application is summarized in conjunction with Figure 19. Referring to Figure 19, taking an electronic device as an example, the method may mainly include some or all of the following contents:
[0283] S1901: Obtain the image L to be processed.
[0284] For example, the image L may be obtained by an electronic device through a screenshot, or may be forwarded by another electronic device. Specific implementation can be found in S1001.
[0285] S1902: Input the image L into a first OCR model for recognition to obtain a first recognition result.
[0286] For specific implementation, please refer to S1002-S1005.
[0287] S1903: Input the image L into the second OCR model for recognition to obtain the image category of the image L.
[0288] For specific implementation, please refer to S1006-S1008.
[0289] S1904: When the image category is an instant messaging chat screenshot, the image L is input into a first edge detection module to obtain attribute information of image and text elements.
[0290] S1905: When the image category is not an instant messaging chat screenshot, the image L is input into a second edge detection module to obtain color block attribute information.
[0291] For example, if the image L is a screenshot of an instant messaging notification card, an order screenshot, or another type of image, the image L is input into the second edge detection module for edge detection.
[0292] For the specific implementation of S1904 and S1905, please refer to S1009-S1011.
[0293] S1906: Determine a filtering rule corresponding to the picture L according to the picture category of the picture L.
[0294] S1907: When the image L is a screenshot of an instant messaging chat, the text recognition result and the edge recognition result of the image L are filtered according to the first filtering rule.
[0295] S1908: When the image L is another screenshot, the text recognition result and the edge recognition result of the image L are filtered according to the second filtering rule.
[0296] For example, other screenshots may be screenshots of instant messaging notification cards, order screenshots, or screenshots of other scenarios.
[0297] As an example of the present application, before filtering according to the second filtering rule, it is also possible to query whether the number of color blocks in the image L is less than the quantity threshold. If the number of color blocks in the image L is less than the quantity threshold, it means that there are not a large number of color blocks in the image L. In this case, the electronic device is usually able to process the image L, so it can be filtered according to the second filtering rule. If the number of color blocks in the image L is greater than or equal to the quantity threshold, it means that the image L includes a large number of color blocks. In this case, filtering may not be performed, but a prompt message may be displayed to guide the user to take a screenshot of the image L again, for example, guiding the user to use the electronic device to capture the area of the image L that includes the schedule information. The quantity threshold can be set according to needs, for example, the quantity threshold can be 10.
[0298] For the specific implementation of S1906 to S1908, please refer to S1012.
[0299] S1909: When the picture L is a photographed picture, output the text recognition result.
[0300] As an example and not a limitation, when picture L is a photographed picture, since the photographed picture may be skewed or include background, filtering processing may not be performed, and the text recognition result may be used for text splicing later.
[0301] In another example, when picture L is a photographed picture, the text recognition result and edge recognition result of picture L can also be filtered according to the second filtering rule, which is not limited in this embodiment of the present application.
[0302] S1910: Based on the filtered edge recognition results and the filtered text recognition results, determine the text line content of each object in the image L, where the object is a filtered chat element or color block.
[0303] For example, when the image L is an instant messaging chat screenshot, the object refers to the filtered chat element, and when the image L is another screenshot, the object refers to the filtered color block. For a specific implementation, see S1013.
[0304] S1911: Determine the scene corresponding to the picture L based on the text line content of each object and the picture category of the picture L.
[0305] For its specific implementation, please refer to S1014.
[0306] S1912: Based on the scene corresponding to the image L, text splicing is performed on the text line recognition content of each object.
[0307] For its specific implementation, please refer to S1015.
[0308] S1913: Construct a prompt based on the scene corresponding to the image L and the spliced text.
[0309] For its specific implementation, please refer to S1016.
[0310] S1914: Recognize the prompt through the NLP model to obtain schedule information.
[0311] For its specific implementation, please refer to S1017-S1019.
[0312] S1915: Perform post-processing of the schedule field on the schedule information.
[0313] Its purpose is to create schedule information. For its specific implementation, please refer to S1020.
[0314] S1916: Displaying schedule information in the calendar application.
[0315] Later, when receiving an edit instruction for the schedule information during the display process, the calendar application displays a smudge interface, which displays content related to the schedule information, such as spliced text. In this way, the user can modify the schedule information through the smudge operation. For specific implementation, please refer to the operation flow in Figure 2.
[0316] In an embodiment of the present application, after receiving a schedule extraction operation for picture L, a first recognition result and a second recognition result are determined by a target recognition model in response to the schedule extraction operation. The first recognition result includes the text recognition result of picture L, and the second recognition result includes the edge recognition result and the picture category of picture L. The edge recognition result includes the graphic element attribute information or color block attribute information of picture L. According to the picture category of picture L, interference information unrelated to the schedule in the text recognition result and the edge recognition result is filtered out, and then based on the filtered first recognition result and the second recognition result, the schedule information of picture L is created and displayed. In this way, the user does not need to manually enter the schedule information item by item in the calendar application, thereby improving the efficiency of creating schedule information.
[0317] FIG20 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Referring to FIG20 , the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. Among them, the sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0318] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0319] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0320] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0321] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0322] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0323] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. The power management module 141 is configured to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.
[0324] The wireless communication functionality of electronic device 100 is implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with a network and other devices via wireless communication technologies.
[0325] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0326] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is an integer greater than one.
[0327] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0328] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0329] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created by the electronic device 100 during use (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0330] The electronic device 100 can implement audio functions, such as music playback and recording, through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D and the application processor.
[0331] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor 180K can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a location different from that of the display screen 194.
[0332] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (such as a coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0333] The above are optional embodiments provided for this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the technical scope disclosed in this application should be included in the scope of protection of this application.
Claims
1. A method for creating schedule information, characterized in that, The method includes: In response to a schedule extraction operation on a first picture, determining a first recognition result and a second recognition result through a target recognition model, where the first recognition result includes a text recognition result of the first picture, and the second recognition result includes an edge recognition result and a picture category of the first picture, and the edge recognition result includes graphic and text element attribute information or color block attribute information of the first picture; Filtering out interference information irrelevant to the schedule in the text recognition result of the first picture and in the edge recognition result according to the picture category of the first picture; Creating schedule information of the first picture based on the filtered first recognition result and the filtered second recognition result; Displaying the schedule information.
2. The method according to claim 1, characterized in that The target recognition model includes a first optical character recognition (OCR) model, a second OCR model, and multiple edge detection models. The first OCR model can be used to determine the text recognition result of a picture, the second OCR model can be used to determine the picture category of a picture, and different edge detection models can be used to perform edge recognition on pictures of different picture categories; The step of, in response to a schedule extraction operation on a first picture, determining a first recognition result and a second recognition result through a target recognition model includes: In response to a schedule extraction operation on the first picture, inputting the first picture into the first OCR model for recognition processing, and outputting the first recognition result; Inputting the first picture into the second OCR model for recognition processing, and outputting the picture category of the first picture; Determining an edge detection model corresponding to the picture category of the first picture from the multiple edge detection models; Inputting the first picture into the determined edge detection model for recognition processing, and outputting the graphic and text element attribute information or color block attribute information of the first picture.
3. The method according to claim 1 or 2, characterized in that, The picture category includes instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. The instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application. The instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application. The order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application program. The other category pictures include other pictures other than the instant messaging chat screenshots, the instant messaging notification card screenshots, and the order screenshots; Wherein, when the first picture is the instant messaging chat screenshot, the edge recognition result includes the graphic and text element attribute information, and when the first picture is one of the instant messaging notification card screenshot, the order screenshot, and the other category pictures, the edge recognition result includes the color block attribute information.
4. The method according to claim 3, characterized in that The step of filtering out interference information irrelevant to the schedule in the text recognition result of the first picture and in the edge recognition result according to the picture category of the first picture includes: Determine the filtering rules corresponding to the picture category of the first picture. Different picture categories correspond to different filtering rules. The instant messaging chat screenshot corresponds to one filtering rule, and the instant messaging notification card screenshot and the order screenshot correspond to the same filtering rule; Filter out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rules corresponding to the picture category of the first picture.
5. The method according to claim 4, characterized in that The first picture is the instant messaging chat screenshot, and the text recognition result includes the text line coordinates; The filtering out of the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rules corresponding to the picture category of the first picture includes: Filter out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture; Filter out the recognition data corresponding to the text lines with a line height less than the target line height in the text recognition result and the edge recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture. The target line height is the average line height of all text lines in the first picture; Filter out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture.
6. The method according to claim 5, wherein The graphic and text element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result also includes the text line recognition content; The filtering out of the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture includes: Determine the chat elements with the category of chat timestamp from the edge recognition result to obtain at least one first candidate chat element; Match the text line recognition content corresponding to each first candidate chat element in the text recognition result of the first picture according to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture; When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element. The target first candidate chat element is any one of the at least one first candidate chat elements; When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filter out the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of the first picture.
7. The method according to claim 6, wherein The determining whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element includes: When it is determined, based on the coordinates of the target first candidate chat element, that the target first candidate chat element is located at the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, the recognition data corresponding to the target first candidate chat element is determined to be filtered; or, When it is determined, based on the coordinates of the target first candidate chat element, that the target first candidate chat element is located at the right position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes time but not date, the recognition data corresponding to the target first candidate chat element is determined to be filtered.
8. The method according to claim 6 or 7, characterized in that, After filtering the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture, it further includes: According to the categories of the remaining chat elements in the edge recognition result, the chat elements irrelevant to the chat content are determined from the remaining chat elements in the edge recognition result to obtain at least one second candidate chat element; According to the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition result of the first picture, the text line recognition content corresponding to each second candidate chat element is matched from the text recognition result of the first picture; The matched text line recognition content and the corresponding text line coordinates are filtered from the text recognition result of the first picture.
9. The method according to claim 4, characterized in that The first picture is a screenshot of the instant messaging notification card, and the text recognition result includes text line coordinates and text line recognition content; Filtering the interference information irrelevant to the schedule from the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture includes: According to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture, the text line recognition content in each color block is matched from the text recognition result of the first picture; For any one of the color blocks, if it is determined according to the text line recognition content in the any one of the color blocks that the any one of the color blocks does not include content related to the schedule, the recognition data corresponding to the any one of the color blocks is filtered from the edge recognition result and the text recognition result of the first picture.
10. The method according to claim 4, characterized in that, The first picture is an order screenshot, the color block attribute information includes color block coordinates, and the text recognition result includes text line coordinates and text line recognition content; Filtering the interference information irrelevant to the schedule from the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture includes: According to the color block coordinates in the edge recognition result and the text line coordinates in the text recognition result of the first picture, the text line recognition content in each color block of the first picture is matched from the text recognition result of the first picture; Determine the color blocks that do not include time and location based on the text line recognition content within each color block; Filter out the recognition data corresponding to the determined color blocks from the text recognition result of the first picture and the edge recognition result.
11. The method according to claim 10, characterized in that, Before matching the text line recognition content within each color block of the first picture from the text recognition result of the first picture according to the color block coordinates in the edge recognition result and the text line coordinates in the text recognition result of the first picture, it further includes: Filter out the recognition data corresponding to the skewed text lines from the text recognition result of the first picture according to the text line coordinates in the text recognition result of the first picture; Filter out the recognition data corresponding to the text lines with a line height less than the target line height from the text recognition result of the first picture and the edge recognition result according to the text line coordinates in the text recognition result of the first picture, where the target line height is the average line height of all text lines in the first picture; Filter out the recognition data corresponding to the middle timestamp from the text recognition result of the first picture and the edge recognition result, where the middle timestamp refers to the text line located in the middle position of the first picture and including only one line of timestamp.
12. The method according to any one of claims 1-11, characterized in that, The text recognition result includes text line coordinates and text line recognition content; Creating the schedule information of the first picture based on the filtered first recognition result and the filtered second recognition result includes: Based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each object from the filtered text recognition result of the first picture, where the object is a chat element or a color block; Determine the scene corresponding to the first picture according to the text line recognition content corresponding to each object and the picture category of the first picture; Perform text splicing on the text line recognition content corresponding to each object based on the scene corresponding to the first picture to obtain spliced text; Construct a prompt based on the spliced text and the scene corresponding to the first picture, where the prompt includes scene description information for describing the scene corresponding to the first picture; Input the prompt into a natural language recognition model for processing to extract the schedule information in the first picture; Create the schedule information.
13. The method according to claim 12, wherein The first picture is an order screenshot; performing text splicing on the text line recognition content corresponding to each object based on the scene corresponding to the first picture to obtain spliced text includes: In the case where the scene corresponding to the first picture is a ticket booking scene, for any one of the remaining color blocks in each remaining color block, if the any one remaining color block includes multiple text lines, perform text splicing line by line; During the splicing process, if there is interactive button text in any one of the remaining color blocks, delete the interactive button text, where the interactive button text is the text line recognition content in the interactive button, the interactive button corresponds to a single text line, the text line length is less than the length threshold, and the text line recognition content is a verb.
14. The method according to claim 12 or 13, characterized in that, After displaying the schedule information, it further includes: In response to an editing operation on the schedule title of the schedule information, a target interface is displayed, and the spliced text is included in the target interface; In response to a selection operation on the content of the text line displayed in the target interface, the text selected by the selection operation is input into the title input box; In response to the end of the editing operation, the schedule title is modified to the content input in the title input box.
15. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of claims 1-14 is implemented.
16. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is caused to execute the method described in any one of claims 1-14.
Citation Information
Patent Citations
Legend identification method and device in engineering drawing, electronic equipment and storage medium
CN113743187A
Information extraction and structuring method for scanned document
CN114299528A
Method, device and equipment for acquiring user interface elements and readable storage medium
CN114627295A
Method, device and system for identifying text information in image
CN116012570A
Text detection method and apparatus, and storage medium
US20190188528A1
Cited By
Multi-dimensional session information extraction method for WeChat chat screenshot
CN121033878A
Interaction method for quickly adding cyclic schedule in waterfall flow type monthly view
CN121807192A