Method for creating schedule information, electronic equipment and readable storage medium

Through the target recognition model, the image is recognized and the agenda information is automatically extracted and created, which solves the problem of inefficient user manual recording of agenda information and realizes efficient and accurate agenda information creation.

CN120258752APending Publication Date: 2025-07-04HONOR DEVICE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311808698.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, users need to manually record the schedule information in third-party social applications in calendar applications, resulting in less efficient creation of schedule information.

Method used

The image is recognized through the target recognition model, filtering out interference information, and automatically extracting and creating schedule information, including schedule information for different picture categories such as instant messaging chat screenshots, notification card screenshots and order screenshots.

Benefits of technology

Improve the efficiency and accuracy of creating agenda information, reduce the steps of manual input by users, and improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258752A_ABST
    Figure CN120258752A_ABST
Patent Text Reader

Abstract

The invention discloses a method for creating schedule information, electronic equipment and a readable storage medium, and belongs to the technical field of terminals. Comprising the steps that a schedule extraction operation on a first picture is responded, a first recognition result and a second recognition result are determined through a target recognition model, the first recognition result comprises a text recognition result of the first picture, and the second recognition result comprises an edge recognition result and a picture category of the first picture; the edge recognition result comprises image-text element attribute information or color block attribute information of the first picture. According to the picture category of the first picture, interference information irrelevant to the schedule in the text recognition result and the edge recognition result of the first picture is filtered out. And based on the filtered first identification result and the filtered second identification result, creating schedule information of the first picture, and displaying the schedule information. The schedule information can be automatically created based on the first picture, manual creation of a user in a calendar application is avoided, and the schedule information creation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and particularly to a method for creating schedule information, an electronic device, and a readable storage medium. Background Art

[0002] With the rapid development of terminal technology, electronic devices can usually install various types of third-party social applications, such as instant messaging applications. Some messages or service notifications in third-party social applications often involve schedule information. For example, the conversation content in the chat window of an instant messaging application involves the schedule information of a certain meeting. In some scenarios, users usually need to record the schedule information involved in third-party social applications through electronic devices.

[0003] In the related art, generally, users need to manually record in the calendar application, which results in low efficiency in creating schedule information. Summary of the Invention

[0004] This application provides a method for creating schedule information, an electronic device, and a readable storage medium, which can solve the problem of low efficiency caused by the need for users to manually create schedule information in the related art. The technical solutions are as follows:

[0005] In a first aspect, a method for creating schedule information is provided. The method includes:

[0006] In response to a schedule extraction operation on a first picture, a target recognition model is used to determine a first recognition result and a second recognition result. The first recognition result includes the text recognition result of the first picture, and the second recognition result includes the edge recognition result and picture category of the first picture. The edge recognition result includes the graphic and text element attribute information or color block attribute information of the first picture. According to the picture category of the first picture, the interference information irrelevant to the schedule in the text recognition result and edge recognition result of the first picture is filtered out. Based on the filtered first recognition result and the filtered second recognition result, the schedule information of the first picture is created and the schedule information is displayed. In this way, it is not necessary for the user to manually create schedule information in the calendar application, improving the efficiency of creating schedule information.

[0007] As an example of this application, the target recognition model includes a first optical character recognition (OCR) model, a second OCR model, and multiple edge detection models. The first OCR model can be used to determine the text recognition result of a picture, the second OCR model can be used to determine the picture category of a picture, and different edge detection models can be used to perform edge recognition on pictures of different picture categories. In response to a schedule extraction operation on a first picture, the specific implementation of determining a first recognition result and a second recognition result through the target recognition model may include: In response to a schedule extraction operation on the first picture, input the first picture into the first OCR model for recognition processing to output a first recognition result. Input the first picture into the second OCR model for recognition processing to output the picture category of the first picture, determine the edge detection model corresponding to the picture category of the first picture from multiple edge detection models, and input the first picture into the determined edge detection model for recognition processing to output the graphic and text element attribute information or color block attribute information of the first picture.

[0008] In this way, after performing text recognition and edge recognition processing on the first picture through the above two branches respectively, a first recognition result and a second recognition result are obtained. Since the text recognition result in the first text recognition result can represent the text content in the first picture, and the edge recognition result in the second recognition result can represent the layout of the first picture, subsequent schedule information extraction based on the two types of data of the first recognition result and the second recognition result can improve the accuracy of information extraction.

[0009] As an example of this application, the picture categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. An instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application. An instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application. An order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application program. Other category pictures include other pictures except instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots.

[0010] Among them, when the first picture is an instant messaging chat screenshot, the edge recognition result includes graphic and text element attribute information. When the first picture is one of an instant messaging notification card screenshot, an order screenshot, and other category pictures, the edge recognition result includes color block attribute information.

[0011] In this way, by classifying the picture categories into the above multiple types, it is possible to adopt a targeted method for picture processing according to the picture category during the subsequent schedule information extraction process, thereby improving the effectiveness and accuracy of schedule information extraction.

[0012] As an example of the present application, according to the picture category of the first picture, the specific implementation of filtering out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture may include: determining the filtering rule corresponding to the picture category of the first picture, where different picture categories correspond to different filtering rules. The instant messaging chat screenshot corresponds to one filtering rule, and the instant messaging notification card screenshot and the order screenshot correspond to the same filtering rule. Filter out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture. In this way, for pictures of different picture categories, different filtering rules are used for filtering, which can specifically filter the interference information in pictures of different picture categories, thereby improving the effectiveness of filtering.

[0013] As an example of the present application, the first picture is an instant messaging chat screenshot, and the text recognition result includes text line coordinates. Accordingly, the specific implementation of filtering out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture may include: filtering out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture. Filter out the recognition data corresponding to the text lines with a line height less than the target line height in the text recognition result and the edge recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture, where the target line height is the average line height of all text lines in the first picture. Filter out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture.

[0014] In this way, by filtering the recognition data corresponding to the skewed lines, small characters, and chat timestamps in the first picture, the interference information unrelated to the schedule can be effectively removed, that is, the information that may interfere with the extraction of schedule information is removed, thereby improving the accuracy of subsequent schedule extraction.

[0015] As an example of the present application, the graphic and text element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result further includes the text line recognition content. Correspondingly, the specific implementation of filtering the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture may include: determining, from the edge recognition result, a chat element whose category is a chat timestamp to obtain at least one first candidate chat element. According to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each first candidate chat element from the text recognition result of the first picture. When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element, and the target first candidate chat element is any one of the at least one first candidate chat element. When it is determined to filter the recognition data corresponding to the target first candidate chat element, filter the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of the first picture.

[0016] In this way, first determine the chat element whose category is a chat timestamp according to the edge recognition result, then match the corresponding text line recognition content from the text recognition result, determine whether it is a chat timestamp through the NLP model according to the matched text line recognition content, and then determine whether it is a chat timestamp according to the position of the chat element, which can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.

[0017] As an example of the present application, the specific implementation of determining whether to filter the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element may include: when it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located in the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, determine to filter the recognition data corresponding to the target first candidate chat element. Or, when it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located on the right side of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes time and does not include date, determine to filter the recognition data corresponding to the target first candidate chat element.

[0018] In this way, judging whether the first candidate chat element is a chat timestamp according to the position characteristics and text line characteristics of the chat timestamp can improve the accuracy of the judgment.

[0019] As an example of the present application, after filtering out the recognition data corresponding to the chat timestamp in the text recognition result and the edge recognition result of the first picture, according to the categories of the remaining chat elements in the edge recognition result, determine the chat elements irrelevant to the chat content from the remaining chat elements in the edge recognition result, to obtain at least one second candidate chat element. According to the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each second candidate chat element from the text recognition result of the first picture. Filter out the matched text line recognition content and the corresponding text line coordinates from the text recognition result of the first picture. In this way, after filtering out the recognition data corresponding to the chat timestamp, small characters, and skewed text lines, further filtering processing is performed on the remaining chat elements to filter out all the interference information in the first picture as much as possible, so as to improve the accuracy of subsequent schedule information extraction.

[0020] As an example of the present application, the first picture is a screenshot of an instant messaging notification card, and the text recognition result includes text line coordinates and text line recognition content. The specific implementation of filtering out the interference information irrelevant to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rules corresponding to the picture category of the first picture may include: according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content in each color block from the text recognition result of the first picture. For any one of the color blocks, if it is determined that the color block does not include content related to the schedule according to the text line recognition content in any one of the color blocks, filter out the recognition data corresponding to any one of the color blocks from the edge recognition result and the text recognition result of the first picture. In this way, by filtering out the color blocks that do not include the schedule, it is convenient to only process the color blocks that include content related to the schedule subsequently, which can improve the data processing efficiency and the accuracy of schedule information extraction.

[0021] As an example of this application, the text recognition result includes the text line coordinates and the recognized content of the text line. Based on the filtered first recognition result and the filtered second recognition result, the specific implementation of creating the schedule information of the first picture may include: Based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the recognized content of the text line corresponding to each object from the text recognition result of the filtered first picture, where the object is a chat element or a color block. Determine the scene corresponding to the first picture according to the recognized content of the text line corresponding to each object and the picture category of the first picture. Based on the scene corresponding to the first picture, perform text splicing on the recognized content of the text line corresponding to each object to obtain the spliced text. Construct a prompt according to the spliced text and the scene corresponding to the first picture, and the prompt includes scene description information, and the scene description information is used to describe the scene corresponding to the first picture. Input the prompt into the natural language recognition model for processing to extract the schedule information in the first picture and create the schedule information. In this way, by adding scene description information to the prompt, the NLP model can extract an accurate schedule information.

[0022] As an example of this application, after the schedule information is displayed, in response to an edit operation on the schedule title of the schedule information, a target interface is displayed, and the target interface includes the spliced text. In response to a selection operation on the text line content displayed in the target interface, the text selected by the selection operation is input into the title input box. In response to the end-of-edit operation, the schedule title is modified to the content input in the title input box. In this way, by displaying the target interface, the user can quickly modify the schedule title by smearing, improving the user experience.

[0023] In a second aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for creating schedule information as described in the first aspect above is implemented.

[0024] In a third aspect, a computer-readable storage medium is provided, and instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is made to execute the method for creating schedule information as described in the first aspect above.

[0025] In a fourth aspect, a computer program product containing instructions is provided. When it runs on a computer, the computer is made to execute the method for creating schedule information as described in the first aspect above.

[0026] The technical effects obtained in the second, third, and fourth aspects above are similar to the technical means obtained in the corresponding first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of an application scenario shown according to an exemplary embodiment;

[0028] Figure 2 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;

[0029] Figure 3 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;

[0030] Figure 4 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;

[0031] Figure 5 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;

[0032] Figure 6 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;

[0033] Figure 7 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;

[0034] Figure 8 It is a schematic diagram of a software system of an electronic device shown according to an exemplary embodiment;

[0035] Figure 9 It is a schematic diagram of an implementation framework for creating schedule information shown according to an exemplary embodiment;

[0036] Figure 10 It is a flowchart of a method for creating schedule information shown according to an exemplary embodiment;

[0037] Figure 11 It is a schematic diagram of the processing of an instant messaging chat screenshot shown according to an exemplary embodiment;

[0038] Figure 12 It is a schematic diagram of the processing of an instant messaging notification card screenshot shown according to an exemplary embodiment;

[0039] Figure 13 It is a flowchart of a method for creating schedule information shown according to another exemplary embodiment;

[0040] Figure 14 It is a schematic diagram of the architecture of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the implementation manners of this application in detail with reference to the accompanying drawings.

[0042] It should be understood that the "multiple" mentioned in this application refers to two or more. In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B. The "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily mean different.

[0043] Referring to "one embodiment" or "some embodiments" described in the specification of this application means that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0044] Before introducing the method for creating schedule information provided by the embodiments of this application, the terms or nouns related to the embodiments of this application will be briefly described first.

[0045] One-stop office software: It can provide a variety of office functions, such as document processing, spreadsheets, presentation production, project management, calendars, and emails. Users can complete multiple office tasks in the same software, improving the work efficiency of users. One-stop office software usually adopts a unified user interface design, making the switching and use between various functional modules more convenient and consistent.

[0046] KV form: It refers to presenting information in the form of a table (it can have a table or not). For example, taking the information presented including two columns of information as an example, the left side is the theme, and the right side is the specific content corresponding to the theme.

[0047] Global collection: It means that users can trigger the electronic device to collect information by swiping three fingers up and down on the screen.

[0048] Magic Text: It is a function for quickly extracting text in pictures. Usually, users can turn on or off the Magic Text switch through the path "Settings > Smart Assistant > Magic Text" to enable or disable this function.

[0049] Prompt: In the embodiments of this application, it refers to the model input data formed by adding a piece of text or instruction to the input text information when using a machine learning model, that is, it includes user prompt and system prompt. The user prompt is the input text information, and the system prompt is a piece of text or instruction added. In this way, it can guide the machine learning model to generate more accurate and targeted outputs. The system Prompt can be a question, a description, a formatted input, or even some keywords. By reasonably designing the prompt, the machine learning model can be guided to understand the intention of the question, so as to generate a relatively accurate answer. In short, Prompt is a technology widely used in machine learning models, which can help users solve data bias, improve the controllability of machine learning models and the problem modeling ability.

[0050] Vertical domain: It refers to providing specific services for a limited group, including industries such as entertainment, medical care, environmental protection, education, and sports.

[0051] Currently, various third-party social applications can be installed in electronic devices. Exemplarily, they include but are not limited to instant messaging (IM) applications, ticketing applications, etc. For example, instant messaging applications can be WeChat, QQ, DingTalk, Feishu, etc., and ticketing applications can be Ctrip, Qunar, Tongcheng, Fliggy, etc. Users can chat and book tickets through third-party social applications, and can also follow the official accounts (or notification accounts) of various industries through third-party social applications to understand the fields and events to be concerned about through the service notification cards issued by the official accounts (or notification accounts). In some scenarios, some messages in third-party social applications may involve schedule information. For example, one or more messages in the chat window of an instant messaging application involve the relevant schedule of an event. Users usually have the need to record schedule information. In this case, generally, users need to manually create schedule information in the calendar application. However, manually creating schedule information is rather cumbersome and the creation efficiency of schedule information is low. Therefore, the embodiments of this application provide a method for creating schedule information, enabling the electronic device to automatically create schedule information and improving the creation efficiency of schedule information.

[0052] As an example of this application, referring to Table 1, the electronic device can create the schedules involved in the following scenarios into schedule information:

[0053] Table 1

[0054]

[0055] The above graphics and text include pictures and text. That is, the method provided in the embodiments of the present application can not only automatically extract schedule information from text, but also automatically extract schedule information from pictures. Specifically, the scope supported by the method provided in the embodiments of the present application is shown in Table 2:

[0056] Table 2

[0057]

[0058] As can be seen from Table 2, the electronic device can extract schedule information from pictures. The pictures can be screenshot pictures or pictures taken by a camera. The form of the screenshot pictures can include, but is not limited to, full-screen screenshots or area screenshots. The windows involved in the screenshot pictures can be full-screen, split-screen or floating windows. The sources of the screenshot pictures usually come from mobile phones, tablets, PCs, etc.; the form of the pictures taken by the camera can be, but is not limited to, printed text, handwritten text and artistic words. In addition, the electronic device can also extract schedule information from text. The form of the text includes, but is not limited to, Text and Webview. The format of the text can be plain text, formatted text or text mixed with pictures. The length of the text can be single-paragraph, multi-paragraph, short text or long text, etc. The embodiments of the present application will be mainly described by taking the input as a picture as an example.

[0059] For ease of understanding, the application scenarios provided in the embodiments of the present application will be introduced next.

[0060] In one example, the multi-round conversations in the chat window of WeChat have a schedule intention. When the user wants to create relevant schedule information, the user can trigger the mobile phone to take a screenshot of the application interface where the chat window is located. For example, the user can double-click on the screenshot trigger area of the mobile phone screen to trigger the screenshot operation, so that the mobile phone takes a screenshot of the interface where the WeChat chat window is located. After that, referring to Figure 1 Figure (a) therein, the mobile phone displays the screenshot editing interface U1 and displays the instant messaging chat screenshot obtained after the screenshot in the screenshot editing interface U1, that is, displays the picture p1. The screenshot editing interface U1 includes a "Share" control. When the user wants to create schedule information in the picture p1, the user can click the "Share" control. Referring to Figure 1 Figure (b) therein, in response to the user's triggering operation on the "Share" control, the mobile phone displays the sharing floating window 10. The sharing floating window 10 includes a calendar icon 11, and the user can click the calendar icon 11. In response to the user's triggering operation on the calendar icon 11, the mobile phone starts to process the picture p1 to extract and create relevant schedule information. Exemplarily, referring to Figure 1In figure (c) during this process, the mobile phone can display a prompt message of "Parsing schedule information offline" so that the user can know that the schedule is being extracted from picture p1 currently. Refer to Figure 1 Figure (d) in it. After the schedule information is successfully created, the mobile phone displays a schedule display interface U2. The schedule display interface U2 includes a schedule display window 12, and the created schedule information 13 is displayed in the schedule display window 12. The schedule information 13 includes content such as a schedule title, time, location, etc. In this way, the mobile phone achieves the purpose of automatically creating and displaying schedule information, which can avoid the need for the user to manually record and improve the efficiency of creating schedule information.

[0061] As an example of this application, the schedule display window 12 also includes multiple editing controls, and the user can also edit the schedule information 13 created by the mobile phone based on the multiple editing controls according to needs. For example, refer to Figure 2 Figure (a) in it. When the user wants to modify the schedule title, they can click on the title editing control 14 in the schedule display window 12. In response to the user's click operation on the title editing control 14, the mobile phone displays a target interface U3 (which can be called a scribbling interface) as shown in Figure 2 Figure (b) in it. Information related to the schedule in picture p1 is displayed in the target interface U3. In this way, the user can scribble on the information displayed in the target interface U3 to modify or fill in the schedule title. Correspondingly, the mobile phone inputs the content scribbled by the user into the title input box 15 in the target interface U3. For example, refer to Figure 2 Figure (b) in it. When the user wants to modify the schedule title to "HarmonyOS Open Discussion", the user can scribble the content of "HarmonyOS", "Open", and "Discussion" in sequence in the target interface U3. Correspondingly, the mobile phone inputs "HarmonyOS", "Open", and "Discussion" into the title input box 15 in sequence. Refer to Figure 2 Figure (c) in it. After the user finishes scribbling, they can trigger the "Input" control (or "√" control) in the target interface U3. In response to the user's triggering operation on the "Input" control (or "√" control), the mobile phone resumes displaying the schedule display window 12. At this time, the user can see from the schedule display window 12 that the schedule title has been modified to the content modified by the user through scribbling operation, that is, the schedule title has been changed from "Open Discussion KaiTan" to "HarmonyOS Open Discussion". In this way, by displaying information related to the schedule and allowing scribbling in the target interface U3, the user can quickly modify the schedule title by scribbling, improving the user experience.

[0062] In addition, refer to Figure 2In Figure (a) thereof, the schedule display window 12 further includes a time editing control. When the user wants to edit the time in the schedule information 13, the user can also modify it based on the time editing control. For example, the user can click on the displayed time to edit it. In addition, after the user swipes down the schedule display window 12, the schedule display window 12 can also provide other editing controls, such as editing controls for the number of repetitions, reminder time, important reminder, etc. Thus, the user can edit the schedule information based on other editing controls, and the embodiments of the present application do not limit this.

[0063] As an example of the present application, after the user clicks on the "√" control in the schedule display window 12, in response to this trigger operation, the mobile phone displays the schedule information 13 in the schedule details area of the calendar application, so that the user can view the schedule information 13 from the schedule details area of the calendar application. As an optional example, after the user clicks on the "√" control in the schedule display window 12, the mobile phone can also display the schedule information 13 in the form of a card at positions such as the desktop, the negative first screen, or the notification center, etc., for the user to quickly view later, and the embodiments of the present application do not limit this.

[0064] It should be noted that the above is only an example of the user triggering the mobile phone to create schedule information through the sharing entry (i.e., the sharing control). In another example, refer to Figure 1 Figure (a) thereof, the mobile phone provides a Magic text control 00 in the screenshot editing interface U1. When the user needs the mobile phone to create schedule information based on the picture p1, the user can also click on the Magic text control 00, thereby triggering the mobile phone to create and display the schedule information with one key.

[0065] In another example, the user can also trigger the mobile phone to create and display schedule information through the Any Door entry. Exemplarily, refer to Figure 3 Figure (a) thereof, after the mobile phone takes a screenshot of the interface where the WeChat chat window is located, the obtained picture p1 is automatically saved to the photo library. Thus, when the user wants the mobile phone to automatically create the schedule information in the picture p1, the user can open the screenshot picture interface U4 in the photo library, and the picture p1 is displayed in the screenshot picture interface U4. Refer to Figure 3 Figure (b) thereof, the user can trigger the mobile phone to select the picture p1, and then, the user can drag the picture p1 to the right side of the mobile phone screen. When the user drags to a certain position, in response to the user's drag operation, the mobile phone displays the application programs that can receive and process the picture p1, such as Figure 3 As shown in Figure (b) thereof, the calendar, WeChat, and QQ are displayed. Thus, the user can continue to drag the picture p1 onto the calendar application and then release it. In response to the user's release operation, the mobile phone starts to process the picture p1 to extract and create relevant schedule information. Refer to Figure 3In figure (c), after the mobile phone creates schedule information, the schedule information is displayed in the schedule display window 12.

[0066] In another example, the user can only select the content related to the schedule in picture p1, and then trigger the mobile phone to extract schedule information from the selected part. For example, see Figure 4 In figure (a), the user can select a part of the content in picture p1. For example, in the screenshot editing interface, controls for triggering the selection operation can be provided. After triggering the control, the user can select on picture p1. In response to the user's selection operation, the mobile phone selects the part selected by the user. As an example, the mobile phone can take a screenshot of the area selected by the user, as shown in Figure 4 figure (b). After that, the user can trigger the mobile phone to extract schedule information from the selected area through interaction entrances such as the "Share" control, and create and display schedule information related to the content in the area. In this way, by supporting the user to select on picture p1, the data processing volume of the mobile phone can be reduced, thereby improving the creation efficiency of schedule information.

[0067] It should be noted that the interaction entrances used by the user to trigger the mobile phone to extract schedule information from picture p1 in the above application scenarios are only exemplary. In some embodiments, the mobile phone can also be triggered to extract schedule information from picture p1 through other interaction entrances. For example, it can also be triggered through interaction entrances such as global collection. The embodiments of the present application do not limit this.

[0068] The above application scenarios are only exemplary. In addition, the mobile phone can also extract schedule information from some types of pictures in other scenarios. Exemplarily, see Figure 5 , there is a forwarded picture p2 in a certain chat window (which can be obtained by screenshot or taken by shooting), and there is schedule information in picture p2. When the user wants to extract the schedule information in picture p2, the user can drag picture p2 to the right side of the mobile phone screen. See Figure 5 In figure (b), when the user drags to a certain position, in response to the user's drag operation, the mobile phone displays application programs that can receive and process picture p2. For example, as shown in Figure 5 figure (b), the calendar, WeChat, and QQ are displayed. After the user drags picture p2 to the calendar application and releases it, in response to the user's release operation, the mobile phone starts to automatically process picture p2 to extract and create relevant schedule information. See Figure 5 In figure (c), after the mobile phone creates schedule information, the schedule information is displayed in the schedule display window 12.

[0069] See Figure 5In figure (c), when the mobile phone creates schedule information multiple times, multiple schedule information can be displayed in the schedule display window 12. When not all schedule information can be fully displayed in the schedule display window 12, the user can trigger the mobile phone to display the hidden schedule information by swiping the schedule display window 12 left and right.

[0070] In addition, the mobile phone not only supports the user to drag the pictures in the chat window to the Any Door entrance, but also supports the user to drag the text to the Any Door entrance. For example, a certain chat window includes chat text, and the chat text includes schedule information. When the user needs the mobile phone to create the schedule information in the chat text, the user can select the chat text and then Figure 5 drag the chat text to the calendar application in the manner shown. Correspondingly, the mobile phone extracts the schedule information in the chat text, creates and displays the relevant schedule information.

[0071] In another example, the mobile phone can also extract schedule information from the screenshot of the instant messaging notification card. For example, refer to Figure 6 figure (a) in. The figure shows a schematic diagram of a screenshot of an instant messaging notification card (i.e., picture p3) shown according to an exemplary embodiment. The picture p3 is obtained by taking a screenshot of the notification card published in the service notification in WeChat, and includes schedule information. When the user wants the mobile phone to create and display the schedule information in the picture p3, the user can trigger the mobile phone to extract the schedule information according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information from the picture p3 through the sharing entrance. Correspondingly, the mobile phone extracts the schedule information based on the picture p3, and then creates or displays the relevant schedule information. For example, refer to Figure 6 figure (b) in. The mobile phone creates and displays schedule information 60, and the schedule information 60 includes the theme, ticket number, departure time and date, location, etc.

[0072] In another example, the mobile phone can also extract schedule information from the order screenshot, where the order screenshot can be a screenshot picture of a hotel order, a train ticket order, an airplane ticket order, etc. For example, refer to Figure 7 figure (a) in. The figure shows a schematic diagram of an order screenshot (i.e., picture p4) shown according to an exemplary embodiment. The picture p4 is obtained by taking a screenshot of the interface where the train ticket order is located, and the picture p4 includes schedule information. When the user wants the mobile phone to create and display the schedule information in the picture p4, the user can trigger the mobile phone according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information from the picture p4 through the Magic Text entrance. Correspondingly, the mobile phone extracts the schedule information based on the picture p4, and then creates or displays the relevant schedule information. For example, the displayed schedule information is as shown in Figure 7 70 in figure (b). The schedule information 70 includes the theme, train number, departure time and arrival time, date, location, etc.

[0073] It should be noted that the above application scenarios are all exemplary and do not limit the application scenarios of the method provided by the embodiments of the present application. In another embodiment, the mobile phone can also extract schedule information from other types of pictures, and other types of pictures include pictures other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots. Of course, other types of pictures can be pictures obtained by taking screenshots or pictures taken by a camera.

[0074] In addition, it should be noted that the above is described by taking the electronic device as a mobile phone as an example. The electronic device involved in the embodiments of the present application can also be an action camera (GoPro), a digital camera, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, etc., and the embodiments of the present application do not make any limitations thereto.

[0075] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. The embodiments of the present application take the Android system with a layered architecture as an example to exemplarily illustrate the software system of the electronic device.

[0076] Figure 8 It is a block diagram of the software system of an electronic device provided by the embodiments of the present application. Refer to Figure 8 , the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system layer, and the kernel layer.

[0077] The application layer may include a series of application packages. Refer to Figure 8 , the application packages may include a calendar and other applications. For example, other applications may include instant messaging, ticket booking, camera, gallery, call, map, navigation, Bluetooth, music, video, short message and other applications.

[0078] In addition, as an example of this application, the application layer further includes a schedule management service and a model management service. The schedule management service can be used to provide services for the calendar, or to be called by the calendar. For example, the schedule management service can create schedule information for the calendar and store data such as schedule information. The schedule management service can include a schedule creation service (which can be called: intelligent parsing processing service) and a schedule database (such as: Calednar Provider schedule database). The schedule management service can create schedule information through the schedule creation service and store the created schedule information through the schedule database. The model management service (which can be called: MagicLive large model service) can be used to provide various models for the schedule management service to call when needed.

[0079] In one example, the model management service can provide a natural language understanding (NLP) model, an object recognition model, and a personal behavior feature model. Among them, the NLP model can be used to identify prompts to determine schedule information. The NLP model can run through a natural language unit (NLU). In some examples, the NLP model can not only identify prompts but also extract keywords in a text, such as extracting keywords like time and location. The object recognition model can be used to perform text recognition and edge recognition on pictures, and can also be used to determine the picture category of the pictures. In one example, the object recognition model includes a first optical character recognition (OCR) model, a second OCR model, and an edge detection model. The first OCR model can be used for text recognition, the second OCR model can be used to determine the picture category, and the edge detection model can be used to perform edge recognition on pictures. The number of edge detection models can be multiple, and different edge detection models can be used to perform edge recognition on pictures of different picture categories; the personal behavior feature model can be used to determine a user profile based on the user's historical behavior data.

[0080] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. As Figure 8 shown, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.

[0081] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core libraries consist of two parts: one part is the functional functions that the Java language needs to call, and the other part is the core libraries of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0082] The system libraries can include multiple functional modules, such as: surface manager, Media Libraries, 3D graphics processing libraries (such as OpenGL ES), 2D graphics engines (such as SGL), etc.

[0083] The kernel layer is the layer between the hardware and the software. The kernel layer includes at least display drivers, camera drivers, audio drivers, and sensor drivers.

[0084] The electronic device can implement the method for creating schedule information provided in the embodiments of the present application through the interaction of the above multiple modules. Exemplarily, see Figure 9 , Figure 9 is a schematic diagram of a method implementation framework shown according to an exemplary embodiment. In implementation, the electronic device can pass the picture to be processed to the calendar application through interaction entrances such as sharing, Any Door, global collection, or Magic Text. Among them, the picture can be a screenshot of the order of train tickets, air tickets, or screenshots of the orders of hotels, catering, entertainment, or screenshots of the orders of sports, health, or screenshots of WeChat mini-programs, web, etc., or screenshots of service notification cards of IM notification numbers, or screenshots of IM chat conversations, etc. By way of example and not limitation, the screenshots of train tickets and air tickets can come from applications such as 12306, China Railway, or Ctrip, and the screenshots of hotel, catering, and entertainment orders can come from applications such as Ctrip, Qunar, Tongcheng, Meituan, Fliggy, Dianping, Damai, etc. The screenshots of sports and health orders can be from applications such as Keep, Lianduoduo, and registration platforms. The screenshots of WeChat mini-programs, web, etc. can include content such as performances, flash sales, and marathons. The screenshots of service notification cards of IM notification numbers can include notification cards released by third parties such as hospitals, scenic spot tickets, educational institutions, and insurance companies through IM notification numbers. The screenshots of IM chat conversations can include content such as work arrangements, invitations, and educational tasks.

[0085] After the calendar application receives a picture, it requests the target recognition model to perform text recognition and edge recognition on the picture. Through text recognition, the text information in the picture can be extracted, and through edge recognition, chat elements or color blocks in the picture can be recognized. For example, see Figure 9 , the picture p1 is passed to the target recognition model through the interaction entry. After text recognition by the target recognition model, the text recognition result is as shown in Figure 9 90 in. In addition, after edge recognition by the target recognition model, the coordinates and categories of each chat element in the picture p1 can be recognized.

[0086] After that, the electronic device can perform filtering processing on the text recognition result and the edge recognition result based on the filtering rules to remove the interference information unrelated to the schedule in the text recognition result and the edge recognition result. As an example of the application, the preset filtering rules may include but are not limited to at least one of the following rules: 1. Discard dense text blocks. 2. Discard small characters. 3. Discard skewed lines. 4. Discard floating text on the drawing. 5. Discard the "back" in the upper left corner of the picture and the "+" symbol in the picture. Exemplarily, after filtering the text recognition result based on the filtering rules, the content shown in Figure 9 91 in can be removed. In addition, after filtering the edge recognition result based on the filtering rules, the coordinates and categories of chat elements unrelated to the schedule can be removed.

[0087] If the text recognition result is not filtered but directly input into the NLP model for schedule information extraction, in this case, referring to Table 3, the recognition success rate of the NLP model is relatively low, usually only reaching about 30%. The reason is the problems of interference information, format line breaks, and loss of layout information. Therefore, in the embodiment of the present application, filtering the text recognition result before extracting schedule information through the NLP model can improve the accuracy of the subsequent NLP model in extracting schedule information.

[0088] Table 3

[0089]

[0090] After the filtering process, the electronic device constructs a prompt (i.e., prompt) based on the remaining text recognition result and edge recognition result after filtering. Then, the prompt is input into the NLP model to extract schedule information through the NLP model. The NLP model outputs schedule information based on the input prompt. For example, Figure 9 92 in shows some content processed by the NLP model. After post-processing the schedule fields of the schedule information, the schedule information to be displayed can be obtained. Exemplarily, the processed schedule information is as shown in Figure 9 93 in. After that, the electronic device can display the schedule information.

[0091] In one example, the post - processing of the schedule field may include, but is not limited to, at least one of the following: 1. Process according to the reminder time rule. 2. Process according to the start time rule of tomorrow. 3. Process according to the details rule. 4. Process based on the title merging pre - filling rule.

[0092] The reminder time rule includes setting the reminder time of the schedule information earlier than the time preset duration extracted by the NLP model. The preset duration can be set according to requirements. For example, if the preset duration is 30 minutes and the time extracted by the NLP model is 8:30, then the reminder time of the schedule information can be set to 8:00. In addition, the reminder time rule also includes repeated reminders. For example, if the schedule information extracted by the NLP model is to grab a number on Monday, Tuesday, and Wednesday, the electronic device sets the reminder time for Monday, the reminder time for Tuesday, and the reminder time for Wednesday for this schedule information, rather than setting only one reminder time.

[0093] The start time rule of tomorrow means that if the time extracted by the NLP model has a situation of crossing days, months, or years, the time after crossing days, months, or years is supplemented. For example, if the date extracted by the NLP model is December 5th, the start time is 23:00, and the end time is 00:30, then the electronic device can supplement the end time as 00:30 on December 6th in the schedule information.

[0094] The details rule means adjusting the layout and font size of the created schedule information according to the size of the screen of the electronic device, so as to be able to display it correctly and clearly in the schedule details area of the calendar application.

[0095] The title merging pre - filling rule includes, in the case where there are multiple different titles corresponding to the same time, selecting the title with the longest length as the title of the schedule information from these multiple titles. In addition, the title merging pre - filling rule also includes using the specified schedule title corresponding to the scenario of the picture as the schedule title of the schedule information. For example, if the scenario of the picture is to make an appointment to get a number for seeing a doctor, and the schedule title extracted by the NLP model is "getting a number" or "taking a number", then the schedule title can be standardized to "registering for seeing a doctor", where the specified schedule titles corresponding to different scenarios can be preset according to requirements.

[0096] It should be noted that the above post - processing of the schedule field is only exemplary. In another example, the post - processing of the schedule field may further include, but is not limited to, at least one of time similarity processing, discarding empty results, risk control, and cleaning non - natural language titles. Time similarity processing means that if the time in the output schedule information is earlier than the current system time of the electronic device, then according to the current system time, the time closest to the time in the schedule information is determined, and the determined time is determined as the time in the schedule information. For example, if the time in the schedule information is Tuesday and the current system time is Wednesday, then it is recorded as Tuesday of the next week in the schedule information. Discarding empty results means discarding the returned empty fields. Risk control means controlling sensitive words. Cleaning non - natural language titles means that if the schedule title does not include a verb, a verb can be added to the schedule title or the schedule title can be default - set according to the scene corresponding to the picture.

[0097] Next, in combination with Figure 10 a detailed introduction to the method for creating schedule information provided by the embodiments of the present application will be given. Refer to Figure 10 and the method may include the following implementation steps:

[0098] S1001: The calendar application receives the picture L to be processed.

[0099] The picture L may be a picture in bitmap format.

[0100] The calendar application receives the picture L passed from the interaction entry. As mentioned above, the interaction entry may be a sharing entry, an arbitrary - door entry, etc. Exemplarily, refer to Figure 1 Figure (a) in, when the user submits the picture L to the calendar application through the sharing control, this interaction entry is the sharing entry.

[0101] S1002: The calendar application sends a schedule creation instruction to the schedule management service, and the picture L is carried in the schedule creation instruction.

[0102] The schedule creation instruction is used to indicate creating and displaying relevant schedule information based on the picture L.

[0103] S1003: The schedule management service sends the picture L to the first OCR model in the model management service.

[0104] In implementation, the schedule management service calls the first OCR model in the model management service and sends the picture L to the first OCR model for text recognition processing.

[0105] S1004: The first OCR model determines the first recognition result of the picture L.

[0106] As an example of the present application, the first recognition result includes the text recognition result of picture L, and the text recognition result includes text block coordinates, text line coordinates, and text line recognition content.

[0107] As an example, the first recognition result further includes target indication information, which can be used to indicate whether the picture L input into the first OCR model is a screenshot picture or a taken picture. Exemplarily, the target indication information can be a first identifier, a second identifier, or a third identifier. The first identifier is used to indicate that the picture L input into the first OCR model is a screenshot picture. The second identifier is used to indicate that the picture L input into the first OCR model is a taken picture and is a picture taken of a document. The third identifier is used to indicate that the picture L input into the first OCR model is other taken pictures, such as pictures taken of advertisements, road signs, magazines, etc. The first identifier, the second identifier, and the third identifier can be set according to requirements. For example, the first identifier is F1, the second identifier is F2, and the third identifier is F3.

[0108] That is, after the first OCR model receives picture L, it recognizes picture L and outputs the first recognition result of picture L.

[0109] S1005: The first OCR model sends the first recognition result to the schedule management service.

[0110] As an example, after receiving the first recognition result, the schedule management service can cache the first recognition result.

[0111] S1006: The schedule management service sends picture L to the second OCR model in the model management service.

[0112] In an example, after receiving the schedule creation instruction sent by the calendar application, in addition to sending picture L to the first OCR model for text recognition processing, the schedule management service can also call the second OCR model in the model management service and send picture L to the second OCR model to determine the picture category through the second OCR model. That is, the operation of S1006 and the operation of S1003 can be executed in parallel.

[0113] Since the method provided in the embodiments of the present application can extract schedule information from pictures of different picture categories, and the content layouts of pictures of different picture categories are different, the electronic device processes pictures of different picture categories in different ways. Therefore, in implementation, after receiving the to-be-processed picture L, the schedule management service not only performs text recognition through the first OCR model, but also inputs picture L into the second OCR model to determine the picture category of picture L.

[0114] As an example of the present application, the picture categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. An instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application; an instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application; an order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application; other category pictures include other pictures other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots, and other category pictures can be screenshot pictures or captured pictures.

[0115] S1007: The second OCR model determines the picture category of picture L.

[0116] After receiving picture L, the second OCR model identifies picture L and outputs the picture category of picture L.

[0117] S1008: The second OCR model sends the picture category of picture L to the schedule management service.

[0118] S1009: The schedule management service sends picture L to the edge detection model corresponding to the picture category of picture L.

[0119] As an example of the present application, multiple edge detection models are provided in the model management service. The multiple edge detection models can be pre-trained, and different edge detection models can perform edge recognition on pictures of different picture categories. In one example, the multiple edge detection models include a first edge detection model and a second edge detection model. The first edge detection model can be used to perform edge recognition on instant messaging chat screenshots to determine the coordinates and categories of chat elements in the instant messaging chat screenshots; the second edge detection model can be used to perform edge recognition on pictures other than instant messaging chat screenshots. For example, the second edge detection model can be used to perform edge recognition on instant messaging notification card screenshots, order screenshots, or other category pictures to determine the coordinates of color blocks in other pictures.

[0120] The different edge detection models can be obtained by iteratively training an initial training model based on picture training samples of pictures of the corresponding picture categories. The picture training samples can be obtained by edge annotation in advance according to requirements. The initial training model can be set according to requirements. As an example but not a limitation, the initial training model can be a recurrent neural network (RNN), etc.

[0121] After receiving the picture category of picture L, the schedule management service determines the edge detection model corresponding to the picture category of picture L from the multiple edge detection models. Exemplarily, when picture L is Figure 1In the case of p1 in Figure (a), that is, in the case of an instant messaging chat screenshot, the edge detection model corresponding to the picture category of picture L is determined from multiple edge detection models as the first edge detection model; when picture L is Figure 6 p3 in Figure (a) in Figure 7 (which is an instant messaging notification card screenshot) or p4 in Figure (a) in

[0122] (which is an order screenshot), the edge detection model corresponding to the picture category of picture L is determined from multiple edge detection models as the second edge detection model. Then, the schedule management service invokes the determined edge detection model and sends picture L to the edge detection model to request edge recognition of picture L.

[0123] The edge recognition result includes graphic and text element attribute information or color block attribute information.

[0124] In one example, when the picture category of picture L is an instant messaging chat screenshot, the edge detection model corresponding to the picture category of picture L is the first edge detection model. After inputting picture L into the first edge detection model for processing, the output edge recognition result includes graphic and text element attribute information. The graphic and text element attribute information includes the attribute information of the chat elements in the instant messaging chat screenshot. The attribute information of the chat elements may include the coordinates and categories of the chat elements. Exemplarily, the chat elements include avatars, titles, nicknames, chat content, chat timestamps, usernames, specified identifiers, etc. The specified identifier includes the "+" identifier. For example, see Figure 11 After inputting picture L into the first edge detection model, the first edge detection model can determine that the chat elements in picture L include Figure 11 multiple items identified by the dashed box in

[0125] Each chat element corresponds to its own coordinates and category. For example, the coordinates of a certain chat element are the coordinates of the four corners of the area where the chat element is located, and the category is an avatar. Optionally, the attribute information of each chat element may further include the serial number of the chat element. The serial numbers of the chat elements in picture L can be set by default according to the sorting of the chat elements in picture L. Figure 12, when the picture L is a screenshot of an instant messaging notification card, after inputting the picture L into the second edge detection model, the second edge detection model can determine that the color blocks in the picture L include Figure 12 multiple items identified by the dashed boxes in , and each color block corresponds to its own coordinates, for example, the coordinates of the four corners of the color block.

[0126] S1011: The edge detection model sends the edge recognition result to the schedule management service.

[0127] As an example, after receiving the edge recognition result, the schedule management service can cache the edge recognition result.

[0128] It is worth mentioning that after performing text recognition and edge recognition processing on the picture L through the above two branches respectively, a first recognition result and a second recognition result can be obtained. The first recognition result includes the text recognition result and the target indication information, and the second recognition result includes the edge recognition result and the picture category. Since the text recognition result can represent the text content in the picture L and the edge recognition result can represent the layout of the picture L, subsequent schedule information extraction based on these two types of data, the first recognition result and the second recognition result, can improve the accuracy of information extraction. The specific implementation can refer to the following steps.

[0129] S1012: When the picture category of the picture L is an instant messaging chat screenshot, the schedule management service filters the text recognition result and the edge recognition result according to the first filtering rule.

[0130] Different picture categories correspond to different filtering rules. In implementation, the schedule management service determines the corresponding filtering rule according to the picture category of the picture L, and then filters the text recognition result and the edge recognition result of the picture L according to the determined filtering rule to filter out the interference information irrelevant to the schedule.

[0131] In an example, the instant messaging chat screenshot corresponds to the first filtering rule, and the first filtering rule may include filtering out the recognition data corresponding to the skewed lines, small characters, and chat timestamps respectively. The skewed lines refer to the text lines corresponding to the chat elements with an inclination angle greater than the preset angle, and the preset angle can be set according to requirements, for example, the preset angle is 10 degrees; the small characters refer to the text lines corresponding to the chat elements with a line height less than the target line height, and the target line height can be the average line height of all text lines in the picture L; the chat timestamp is a timestamp used to indicate the chat time. Optionally, the first filtering rule may also include filtering out dense text blocks, floating text on the picture, "back" in the upper left corner, "+", etc. in the lower right corner. Exemplarily, see Figure 11 , Figure 11Identify the partial chat elements in picture L that need to be filtered according to an exemplary embodiment, including chat timestamp 1101, small text 1102, skewed line 1103, and "+" 1004.

[0132] As an example, in the implementation of filtering the recognition data corresponding to the skewed line, the inclination angle of each text line can be calculated according to the text line coordinates in the text recognition result, so as to determine which text lines are skewed lines, and then delete the recognition data corresponding to the skewed lines. For example, delete the text line coordinates and text line recognition content of the skewed lines.

[0133] As an example, in the implementation of filtering the recognition data corresponding to the small text, the line height of each text line in picture L can be determined according to the text line coordinates in the text recognition result, so as to filter the recognition data corresponding to the text lines with a line height less than the target line height. For example, filter the text line coordinates and text line recognition content of the text lines with a line height less than the target line height.

[0134] As an example, in the implementation of filtering the recognition data corresponding to the chat timestamp, the chat elements whose category is the chat timestamp can be determined according to the edge recognition result, and at least one first candidate chat element can be obtained. According to the coordinates of each first candidate chat element in at least one first candidate chat element and the text line coordinates in the text recognition result of picture L, match the text line recognition content corresponding to each first candidate chat element in the text recognition result of picture L. When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element. The target first candidate chat element is any one of at least one first candidate chat element. When it is determined to filter the recognition data corresponding to the target first candidate chat element, filter the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of picture L.

[0135] As described above, the edge recognition result of the instant messaging chat screenshot includes the categories of chat elements. Therefore, according to the categories of chat elements in the edge recognition result, chat elements with the category of chat timestamp can be filtered out to obtain at least one first candidate chat element. Since the edge recognition result does not include the text line recognition content, that is, it is impossible to know the text content corresponding to each chat element. In some possible cases, the edge detection model may misjudge the category of a chat element that is not a chat timestamp as a chat timestamp. Therefore, in order to avoid filtering out chat elements that are not chat timestamps as much as possible, after determining at least one first candidate chat element, the text line recognition content corresponding to each first candidate chat element can be matched from the text recognition result of picture L according to the coordinates of each first candidate chat element and the text line coordinates of picture L. For example, for any one of the first candidate chat elements, according to the coordinates of this first candidate chat element and the text line coordinates of picture L, the text line recognition content of at least one text line located in the area corresponding to this first candidate chat element is determined from the text recognition result of picture L, so as to match the text line recognition content corresponding to this first candidate chat element. After that, the schedule management service can determine whether the matched text line recognition content is a timestamp through the NLP model. For example, the schedule management service can send the matched text line recognition content to the NLP model through getEntitiy() to request the NLP model to identify whether the text line recognition content is a timestamp. If it is determined through the NLP model that the text line recognition content is a timestamp, it can be further determined that this first candidate chat element may be a chat timestamp. Otherwise, if it is determined through the NLP model that the text line recognition content is not a timestamp, it can be determined that this first candidate chat element is not a chat timestamp.

[0136] Since the position of the chat timestamp in the instant messaging chat screenshot is generally fixed. For example, it is usually centered in the chat window of WeChat. Therefore, in the case where it is determined through the NLP model that a certain first candidate chat element may be a chat timestamp, it can be further determined whether this first candidate chat element is a chat timestamp according to the coordinates of this first candidate chat element, so as to improve the accuracy of chat timestamp recognition.

[0137] As an example of the present application, for any one of the at least one first candidate chat elements, the specific implementation of determining whether to filter out the recognition data corresponding to this first candidate chat element according to the coordinates of this first candidate chat element may include the following two cases:

[0138] The first case: When it is determined according to the coordinates of this first candidate chat element that this first candidate chat element is located in the middle position of Picture L and the area corresponding to this first candidate chat element includes a single text line, determine to filter out the recognition data corresponding to this first candidate chat element.

[0139] Since in the chat windows of most instant messaging applications, the chat timestamp is generally centered and the text line recognition content in the chat timestamp only includes one line, that is, only includes the timestamp. Therefore, if it is determined that this first candidate chat element is located in the middle position of Picture L and the area corresponding to this first candidate chat element only includes a single text line, and since it has been determined through the NLP model that the text line recognition content is the timestamp, it can be determined that this first candidate chat element is the chat timestamp, that is, it is determined that the recognition data corresponding to this first candidate chat element can be filtered out.

[0140] The second case: When it is determined according to the coordinates of this first candidate chat element that this first candidate chat element is located on the right side of Picture L and the area corresponding to this first candidate chat element includes a single text line, if the text line recognition content corresponding to this first candidate chat element only includes time and does not include date, determine to filter out the recognition data corresponding to this first candidate chat element.

[0141] Since in the chat windows of some instant messaging applications (such as the chat interface forwarded in WeChat), the chat timestamp may also be displayed on the right side, and the text line recognition content in the chat timestamp only includes one line, and in addition, the chat timestamp only includes time and does not include date. Therefore, if it is determined that this first candidate chat element is located on the right side of Picture L and the area corresponding to this first candidate chat element only includes a single text line, it can be judged whether the text line recognition content corresponding to this first candidate chat element includes a date. For example, it can be determined through the NLP model whether the text line recognition content includes a date. If it is determined that it does not include a date, it can be determined that this first candidate chat element is the chat timestamp, that is, it is determined that the recognition data corresponding to this first candidate chat element can be filtered out. Of course, if it is determined that it includes a date, it can be determined not to filter.

[0142] In an example, for the second case, it is also possible not to judge whether the text line recognition content corresponding to this first candidate chat element only includes time and does not include date. As long as it is determined that this first candidate chat element is located on the right side of Picture L and the area corresponding to this first candidate chat element includes a single text line, the schedule management service can determine to filter out the recognition data corresponding to this first candidate chat element.

[0143] When it is determined through the above process that a certain first candidate chat element is a chat timestamp, the schedule management service deletes the recognition data corresponding to this first candidate chat element from the text recognition result and the edge recognition result of picture L. For example, it deletes the text line coordinates and the recognized text content corresponding to this first candidate chat element from the text recognition result of picture L, and deletes the coordinates and category corresponding to this first candidate chat element from the edge recognition result of picture L. Of course, if it is determined through the above process that a certain first candidate chat element is not a chat timestamp, the schedule management service does not filter out the recognition data corresponding to this first candidate chat element.

[0144] It is worth mentioning that first determining the chat elements whose category is a chat timestamp according to the edge recognition result, then matching the corresponding recognized text content from the text recognition result, determining whether it is a chat timestamp through the NLP model based on the matched recognized text content, and then determining whether it is a chat timestamp according to the position of the chat element can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.

[0145] It should be noted that the above first filtering rule is only exemplary. When the instant messaging chat screenshots are from different instant messaging applications, the layout of their chat elements is usually different, and the chat elements may also be different, so that the interference information in the instant messaging chat screenshots of different instant messaging applications may be different. Exemplarily, it usually includes several possible situations shown in Table 4. Therefore, in another example, the first filtering rule may also include other rules for filtering out information irrelevant to the chat content.

[0146] Table 4

[0147]

[0148] In order to effectively filter out the interference information in the instant messaging chat screenshots from different instant messaging applications, a first filtering rule can be set according to the union of the possible interference information shown in Table 4, so as to ensure that no matter which instant messaging chat screenshot is processed, the interference information can be effectively removed. Exemplarily, the first filtering rule can also include filtering out the text in the avatar, the user name, the specified identifier, etc. In implementation, after filtering out the recognition data corresponding to the chat timestamp in the picture L from the text recognition result and the edge recognition result of the picture L, the schedule management service can determine the chat elements irrelevant to the chat content from the remaining chat elements in the edge recognition result according to the category of each remaining chat element in the edge recognition result, to obtain at least one second candidate chat element, and match the text line recognition content corresponding to each second candidate chat element in the text recognition result of the picture L according to the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition result of the picture L. Delete the matched text line recognition content and the corresponding text line coordinates from the text recognition result of the picture L.

[0149] Exemplarily, in the implementation of filtering out the text in the avatar, the schedule management service can determine the chat element whose category is the avatar according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the text recognition result of the picture L according to the coordinates of the chat element, and if there is a matched text line recognition content, delete the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content from the text line recognition result of the picture L, so as to delete the text in the avatar.

[0150] Exemplarily, in the implementation of filtering out the user name, the schedule management service can determine the chat element whose category is the user name according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the text recognition result of the picture L according to the coordinates of the chat element, and if there is a matched text line recognition content, delete the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content from the text line recognition result of the picture L, so as to delete the user name.

[0151] Exemplarily, in the implementation of filtering out the specified identifier, the schedule management service can determine the chat element whose category is the specified identifier according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the edge recognition result of the picture L, and if the text line recognition content is the specified identifier, for example, it is "+", the recognition data corresponding to the chat element can be deleted from the text line recognition result of the picture L, for example, the text line coordinates and the text line recognition content corresponding to the chat element are deleted. The schedule management service can also delete the recognition data corresponding to the chat element from the edge recognition result, for example, delete the coordinates and the category of the chat element.

[0152] It should be noted that the above example is based on the case where picture L is a screenshot of an instant messaging chat. In another example, if picture L is not a screenshot of an instant messaging chat, such as a screenshot of an instant messaging notification card or an order screenshot, the schedule management service filters out interference information based on the second filtering rule. In one example, in the implementation of filtering based on the second filtering rule, the schedule management service can match the text line recognition content in each color block from the text recognition result of picture L according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of picture L. For any one of the color blocks, if it is determined that the content related to the schedule is not included in any one of the color blocks according to the text line recognition content in any one of the color blocks, such as not including time and location, the recognition data corresponding to any one of the color blocks is filtered out from the edge recognition result and the text recognition result of picture L. In one example, an NLP model can be used to determine whether time and location are included in the color block. For example, the text line recognition content in the color block can be sent to the NLP model one by one to request the NLP model to determine whether time and location are included.

[0153] In the case where picture L is a screenshot of an instant messaging notification card, it can be seen from Table 4 that its possible interference information includes small characters, skewed lines, and a middle timestamp. Therefore, in one example, before filtering based on the second filtering rule, the recognition data corresponding to the skewed lines, small characters, and the middle timestamp can be filtered out from the text line recognition content of picture L and the edge recognition result first, and then filtered out based on the second filtering rule. When filtering the middle timestamp, it can be determined whether it is the middle timestamp by matching the corresponding text line recognition content and combining the coordinates of the color block according to the text line recognition content.

[0154] As an example of this application, before filtering, the schedule management service can also determine whether there are dense text blocks according to the text block coordinates in the text recognition result of picture L. If there are dense text blocks, a prompt message can be displayed in the calendar application. The prompt message is used to prompt the user that there are dense text blocks, so that the user can re-crop picture L according to the needs. If there are no dense text blocks, the schedule management service performs the filtering operation.

[0155] Or in another example, during the filtering process, if the schedule management service determines that there are dense text blocks according to the text block coordinates in the text recognition result of picture L, the text line coordinates and text line recognition content in these text blocks are deleted.

[0156] It is worth mentioning that if no filtering is performed, it is likely to cause the subsequent NLP model to be unable to accurately extract schedule information. For example, for Figure 9In the picture p1, it may extract the time as the chat timestamp "20:37", and may extract the address as "No. 5, Danling Street, Ningyuan Road, Haidian District, Beijing", etc. In the embodiments of the present application, after determining the first recognition result and the second recognition result, filtering out the interference information in the picture L can improve the accuracy of subsequent schedule information extraction.

[0157] S1013: The schedule management service matches the text line recognition content of each remaining chat element from the filtered text recognition result based on the filtered edge recognition result.

[0158] The schedule management service matches the text line recognition content corresponding to each filtered chat element from the filtered text recognition result according to the coordinates of each chat element in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.

[0159] When the picture L is not an instant messaging chat screenshot, the schedule management service matches the text line recognition content in each filtered color block from the filtered text recognition result according to the coordinates of each color block in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.

[0160] S1014: The schedule management service determines the scene corresponding to the picture L according to the text line recognition content of each remaining chat element and the picture category of the picture L.

[0161] The scene corresponding to the picture L refers to the scene involved in the text content in the picture L.

[0162] When it is determined that the picture L is an instant messaging chat screenshot according to the picture category of the picture L, if it is further determined to be a chat conversation according to the text line recognition content of each filtered chat element, the scene corresponding to the picture L can be determined as the IM chat scene.

[0163] In addition, when the picture category of the picture L is other pictures, such as an instant messaging notification card screenshot or an order screenshot, the schedule management service can also combine the text line recognition content of each matched color block to determine the scene corresponding to the picture L, such as the service notification card scene or the high-speed rail travel order scene.

[0164] S1015: The schedule management service performs chat conversation splicing on the text line recognition content of each filtered chat element based on the scene corresponding to the picture L.

[0165] In implementation, the schedule management service performs operations such as line breaks, carriage returns, and splicing on the text line recognition content of each filtered chat element according to the scene corresponding to Picture L. Since chat content is usually rather casual, for example, a complete sentence may be sent in multiple messages. When it is determined that the scene corresponding to Picture L is an IM chat scene, the schedule management service can splice these multiple messages into one sentence. Therefore, performing chat conversation splicing based on the scene corresponding to Picture L can make the spliced text closer to natural language text. By way of example and not limitation, the text after chat conversation splicing is consistent with the content displayed in the smearing interface.

[0166] In one example, the content of the spliced chat conversation includes the conversation type, title, nickname, conversation content, etc., and the nickname can be customized. For example, the chat conversation content can be spliced according to the following format:

[0167] Conversation topic: IM chat

[0168] Title: xx group chat

[0169] Nickname A: xxx

[0170] Nickname B: xxx

[0171] Nickname A: xxxx .......

[0173] It should be noted that S1014 to S1016 are optional operations. In another example, the schedule management service can also perform chat conversation splicing on the text line recognition content of each filtered chat element according to the picture category of Picture L.

[0174] In addition, when Picture L is other pictures, such as a screenshot of an instant messaging notification card or an order screenshot, the schedule management service splices the text line recognition content in each color block after filtering according to the scene corresponding to Picture L, and its implementation can refer to the splicing of chat conversations.

[0175] S1016: The schedule management service constructs a prompt according to the scene corresponding to Picture L and the spliced text.

[0176] Pictures of different picture categories correspond to different prompt construction templates. When constructing a prompt, the schedule management service can use the prompt construction template corresponding to the picture category of Picture L to construct the prompt. As an example of this application, the constructed prompt includes scene description information, and the scene description information is used to indicate the scene corresponding to Picture L, that is, when constructing the prompt, the schedule management service adds the scene description information of Picture L.

[0177] It should be noted that if scene description information is not added when constructing the prompt, then when the NLP model performs recognition later, the NLP model is likely to extract each time and the verbs related to before and after that time into a schedule information. In this way, multiple schedule information are easily extracted, resulting in inaccurate extraction of schedule information. Therefore, in order to enable the NLP model to accurately identify schedule information, the schedule management service adds scene description information to the constructed prompt, so that the NLP model can extract an accurate schedule information.

[0178] Exemplarily, taking the picture L as an instant messaging chat screenshot, the prompt constructed by the schedule management service can be:

[0179] <|Human|>The following content may be an IM chat conversation. The IM chat conversation contains <title, time, location, participants> fields. Each event's existing fields are output in one line in json format without additional reply. \n\nKaiTan

[0180] HarmonyOS Architecture Evolution and Key Technologies

[0181] HDC Together

[0182] Time 09:00 - 16:30, October 23rd

[0183] Location Microsoft Asia-Pacific R & D Group Building

[0184] No. 5, Danling Street, Haidian District, Beijing

[0185] <|Moss|>

[0186] S1017: The schedule management service sends the prompt to the NLP model.

[0187] After the schedule management service constructs the prompt, it calls the NLP model and sends the constructed prompt to the NLP model for recognition to extract schedule information. Exemplarily, the schedule management service can request the NLP model to extract schedule information through extractinformation().

[0188] As described above, the NLP model not only has the ability to recognize schedule information but also has the ability to extract keywords. Since the NLP model may have a certain time delay in recognizing the prompt, if the recognition duration is long, it will affect the user experience. Therefore, in some examples, the schedule management service can also send the spliced text to the NLP model line by line at the behavior granularity to request the NLP model to extract the keywords in the spliced text. In this way, in the case of a long recognition duration, the extracted keywords can be used to construct the schedule information. Exemplarily, the schedule management service can instruct the NLP model to extract keywords through getEntity(). For example, the NLP model can be specified to extract time and location keywords, that is, the module is specified as time and location.

[0189] It should be noted that the embodiments of the present application are described by taking the NLP model deployed in the electronic device as an example. In another example, the NLP model can also be deployed in the cloud, and the cloud can provide an interface for the electronic device to call the NLP model. In this way, when the NLP model is needed, the schedule management service can call the NLP model through the provided interface. The embodiments of the present application do not limit this.

[0190] S1018: The NLP model determines the schedule information based on the prompt.

[0191] In one example, the schedule information output by the NLP model includes: {"data": "[{'title': 'KaiTan', 'time','start time': '09:00', 'end time': '16:30', 'date': 'October 23rd', 'location': 'No. 5, Danling Street, Haidian District, Beijing, Microsoft Asia-Pacific R & D Group Building'}]"}.

[0192] In one example, when the schedule management service also sends the spliced text to the NLP model to instruct the extraction of keywords, the schedule management service can also extract the corresponding keywords.

[0193] S1019: The NLP model sends the schedule information to the schedule management service.

[0194] Exemplarily, the NLP model can send the schedule information to the schedule management service in JSON format.

[0195] In addition, if the schedule management service extracts keywords, it also sends the extracted keywords to the schedule management service.

[0196] S1020: The schedule management service performs post-processing on the schedule fields of the schedule information.

[0197] The post-processing of the schedule fields of the schedule information can be referred to above.

[0198] In one example, if the NLP model returns schedule information within a specified duration, the schedule management service performs post-processing on the schedule fields based on the schedule information fed back by the NLP model to create schedule information. Since the schedule information is recognized based on the prompt, the accuracy of schedule information creation can be improved. If the NLP model does not return schedule information within the specified duration, it indicates that the schedule information recognition has timed out. In this case, the schedule management service can construct schedule information based on the keywords extracted by the NLP model to avoid, as much as possible, the problem of lag in schedule information creation. Among them, the specified duration is set according to requirements. For example, the specified duration can be 5s.

[0199] It should be noted that when creating schedule information based on keywords, post-processing of schedule fields can also be performed and then created according to a preset template, such as including content like the subject, time, location, etc.

[0200] In one example, before creating schedule information, the schedule management service can also call the personal behavior feature model to request and query the user's historical behavior data. For example, the historical behavior data includes historical locations, etc. Correspondingly, the personal behavior feature model returns the historical behavior data. In this way, the schedule management service can predict the places the user may go based on the historical behavior data, and then combine the schedule information fed back by the NLP model to create the final schedule information. For example, add the predicted address information to the schedule information.

[0201] S1021: The schedule management service displays the processed schedule information in the calendar application.

[0202] Exemplarily, when the picture L is an instant messaging chat screenshot, the schedule information displayed on the electronic device is as shown in Figure 1 13 in figure (d) as shown.

[0203] In one example, before displaying the schedule information, a request confirmation notice can also be displayed first. After receiving the confirmation display instruction triggered by the user based on the request confirmation notice, the schedule management service then displays the schedule information in the calendar application.

[0204] As an example of this application, the electronic device also supports the user to edit the displayed schedule information. Exemplarily, it supports the user to modify the schedule title of the schedule information. For example, see Figure 2The embodiments shown. In this process, when the user triggers the electronic device to display the target interface, it is necessary to display the text that can be smeared in the target interface. For this purpose, the schedule management service can also request the NLP model to perform word combination processing on the spliced text. For example, combine the two words "I" and "men" into "we" so that the text can be displayed in the target interface according to the combined words, thus facilitating the user to smear. In implementation, the schedule management service can call the NLP model through getWordSegment() to request the NLP model to perform word combination processing. In addition, the schedule management service can also call the NLP model through getWordSegment() to request the NLP model to determine the theme of the spliced text. For example, specify the NLP model to extract the theme entities (such as meetings, dinners). In this way, the schedule management service can display the theme in the target interface.

[0205] In the embodiments of the present application, after receiving the schedule extraction operation for Picture L, in response to the schedule extraction operation, the first recognition result and the second recognition result are determined through the target recognition model. The first recognition result includes the text recognition result of Picture L, and the second recognition result includes the edge recognition result and the picture category of Picture L. The edge recognition result includes the graphic element attribute information or the color block attribute information of Picture L. Filter out the interference information irrelevant to the schedule in the text recognition result and the edge recognition result according to the picture category of Picture L, and then create and display the schedule information of Picture L based on the filtered first recognition result and second recognition result. In this way, it is not necessary for the user to manually input the schedule information item by item in the calendar application, improving the efficiency of creating schedule information.

[0206] Next, in combination with Figure 13 a general introduction to the implementation process of the method for creating schedule information provided in the embodiments of the present application will be given. See Figure 13 Taking the electronic device as the execution subject as an example, this method mainly may include the following parts or all of the content:

[0207] S1301: Obtain the picture L to be processed.

[0208] For example, Picture L can be obtained by the electronic device through screenshot, or it can also be forwarded by other electronic devices. The specific implementation can refer to S1001.

[0209] S1302: Input Picture L into the first OCR model for recognition to obtain the first recognition result.

[0210] The specific implementation can refer to S1002 - S1005.

[0211] S1303: Input Picture L into the second OCR model for recognition to obtain the picture category of Picture L.

[0212] For specific implementation, please refer to S1006-S1008.

[0213] S1304: When the image category is an instant messaging chat screenshot, the image L is input into a first edge detection module to obtain attribute information of image and text elements.

[0214] S1305: When the image category is not an instant messaging chat screenshot, the image L is input into a second edge detection module to obtain color block attribute information.

[0215] For example, if the image L is a screenshot of an instant messaging notification card, an order screenshot, or another type of image, the image L is input into the second edge detection module for edge detection.

[0216] For the specific implementation of S1304 and S1305, please refer to S1009-S1011.

[0217] S1306: Determine a filtering rule corresponding to the picture L according to the picture category of the picture L.

[0218] S1307: When the image L is a screenshot of an instant messaging chat, the text recognition result and the edge recognition result of the image L are filtered according to the first filtering rule.

[0219] S1308: When the picture L is another screenshot, the text recognition result and the edge recognition result of the picture L are filtered according to the second filtering rule.

[0220] For example, other screenshots may be screenshots of instant messaging notification cards, order screenshots, or screenshots of other scenarios.

[0221] As an example of the present application, before filtering according to the second filtering rule, it is also possible to query whether the number of color blocks in the image L is less than the number threshold. If the number of color blocks in the image L is less than the number threshold, it means that there are not a large number of color blocks in the image L. In this case, the electronic device is usually able to process the image L, so the filtering process can be performed according to the second filtering rule. If the number of color blocks in the image L is greater than or equal to the number threshold, it means that the image L includes a large number of color blocks. In this case, filtering may not be performed, but the user may be guided to re-screenshot the image L by displaying a prompt message, such as guiding the user to capture a portion of the area in the image L including the schedule information through the electronic device. The number threshold can be set according to demand, for example, the number threshold can be 10.

[0222] For the specific implementation of S1306 to S1308, please refer to S1012.

[0223] S1309: When the picture L is a photographed picture, output the text recognition result.

[0224] As an example rather than a limitation, when Picture L is a captured picture, since the captured picture may be skewed or include a background, etc., filtering processing may not be performed, and the text recognition results are subsequently used for text splicing.

[0225] In another example, when Picture L is a captured picture, the text recognition results and edge recognition results of Picture L may also be filtered according to the second filtering rule, and the embodiments of the present application do not limit this.

[0226] S1310: Based on the filtered edge recognition results and the filtered text recognition results, determine the text line content of each object in Picture L, where the object is the filtered chat element or color block.

[0227] Exemplarily, when Picture L is an instant messaging chat screenshot, the object refers to the filtered chat element; when Picture L is other screenshots, the object refers to the filtered color block. The specific implementation can refer to S1013.

[0228] S1311: Determine the scene corresponding to Picture L according to the text line content of each object and the picture category of Picture L.

[0229] The specific implementation can refer to S1014.

[0230] S1312: Based on the scene corresponding to Picture L, perform text splicing on the text line recognition content of each object.

[0231] The specific implementation can refer to S1015.

[0232] S1313: Based on the scene corresponding to Picture L and the spliced text, construct a prompt message.

[0233] The specific implementation can refer to S1016.

[0234] S1314: Use the NLP model to recognize the prompt message to obtain schedule information.

[0235] The specific implementation can refer to S1017 - S1019.

[0236] S1315: Perform post - processing on the schedule fields of the schedule information.

[0237] The purpose is to create schedule information. The specific implementation can refer to S1020.

[0238] S1316: Display the schedule information in the calendar application.

[0239] After that, during the process of displaying schedule information, when an editing instruction for the schedule information is received, the calendar application displays a smearing interface, and the content related to the schedule information is displayed in the smearing interface, such as the displayed spliced text. In this way, the user can modify the schedule information through the smearing operation. For the specific implementation, reference can be made to Figure 2 the operation process.

[0240] In the embodiment of the present application, after receiving the schedule extraction operation for the picture L, in response to the schedule extraction operation, the first recognition result and the second recognition result are determined through the target recognition model. The first recognition result includes the text recognition result of the picture L, and the second recognition result includes the edge recognition result and the picture category of the picture L. The edge recognition result includes the graphic element attribute information or the color block attribute information of the picture L. The interference information irrelevant to the schedule in the text recognition result and the edge recognition result is filtered according to the picture category of the picture L, and then based on the filtered first recognition result and the second recognition result, the schedule information of the picture L is created and displayed. In this way, it is not necessary for the user to manually input the schedule information item by item in the calendar application, which improves the efficiency of creating the schedule information.

[0241] Figure 14 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Refer to Figure 14 , the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0242] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0243] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0244] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0245] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0246] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods in the above embodiments.

[0247] The charging management module 140 is used to receive charging input from a charger. Here, the charger can be a wireless charger or a wired charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, the wireless communication module 160, etc.

[0248] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.

[0249] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0250] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is an integer greater than 1.

[0251] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0252] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0253] The internal memory 121 can be used to store computer-executable program codes, and the computer-executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0254] The electronic device 100 can implement audio functions, such as music playback, recording, etc., through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor.

[0255] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from the display screen 194.

[0256] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.

[0257] The above are the optional embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the technical scope disclosed in the present application shall be included within the protection scope of the present application.

Claims

1. A method for creating schedule information, characterized in that, The method includes: In response to a schedule extraction operation on a first picture, determining a first recognition result and a second recognition result through a target recognition model, where the first recognition result includes a text recognition result of the first picture, and the second recognition result includes an edge recognition result and a picture category of the first picture, and the edge recognition result includes graphic and text element attribute information or color block attribute information of the first picture; Filtering out interference information unrelated to the schedule in the text recognition result of the first picture and the edge recognition result according to the picture category of the first picture; Creating schedule information of the first picture based on the filtered first recognition result and the filtered second recognition result; Displaying the schedule information.

2. The method according to claim 1, characterized in that The target recognition model includes a first optical character recognition (OCR) model, a second OCR model, and multiple edge detection models. The first OCR model can be used to determine the text recognition result of a picture, the second OCR model can be used to determine the picture category of a picture, and different edge detection models can be used to perform edge recognition on pictures of different picture categories; The step of, in response to a schedule extraction operation on a first picture, determining a first recognition result and a second recognition result through a target recognition model includes: In response to a schedule extraction operation on the first picture, inputting the first picture into the first OCR model for recognition processing, and outputting the first recognition result; Inputting the first picture into the second OCR model for recognition processing, and outputting the picture category of the first picture; Determining an edge detection model corresponding to the picture category of the first picture from the multiple edge detection models; Inputting the first picture into the determined edge detection model for recognition processing, and outputting the graphic and text element attribute information or color block attribute information of the first picture.

3. The method according to claim 1 or 2, characterized in that The picture category includes instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. The instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application. The instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application. The order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application program. The other category pictures include other pictures except the instant messaging chat screenshots, the instant messaging notification card screenshots, and the order screenshots; Wherein, when the first picture is an instant messaging chat screenshot, the edge recognition result includes the graphic and text element attribute information, and when the first picture is one of the instant messaging notification card screenshot, the order screenshot, and the other category pictures, the edge recognition result includes the color block attribute information.

4. The method according to claim 3, characterized in that, The step of filtering out interference information unrelated to the schedule in the text recognition result of the first picture and the edge recognition result according to the picture category of the first picture includes: Determine the filtering rules corresponding to the picture category of the first picture. Different picture categories correspond to different filtering rules. The instant messaging chat screenshot corresponds to one filtering rule, and the instant messaging notification card screenshot and the order screenshot correspond to the same filtering rule; Filter out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture.

5. The method according to claim 4, characterized in that The first picture is the instant messaging chat screenshot, and the text recognition result includes the text line coordinates; The filtering out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture includes: Filter out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture; Filter out the recognition data corresponding to the text lines with a line height less than the target line height in the text recognition result and the edge recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture. The target line height is the average line height of all text lines in the first picture; Filter out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture.

6. The method according to claim 5, characterized in that, The graphic and text element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result also includes the text line recognition content; The filtering out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture includes: Determine the chat elements with the category of chat timestamp from the edge recognition result to obtain at least one first candidate chat element; Match the text line recognition content corresponding to each first candidate chat element in the text recognition result of the first picture according to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture; When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element. The target first candidate chat element is any one of the at least one first candidate chat elements; When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filter out the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of the first picture.

7. The method according to claim 6, wherein The determining whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element includes: When it is determined, according to the coordinates of the target first candidate chat element, that the target first candidate chat element is located in the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, the recognition data corresponding to the target first candidate chat element is determined to be filtered; or, When it is determined, according to the coordinates of the target first candidate chat element, that the target first candidate chat element is located on the right side of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes time and does not include date, the recognition data corresponding to the target first candidate chat element is determined to be filtered.

8. The method according to claim 6 or 7, characterized in that After filtering the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture, it further includes: According to the categories of the remaining chat elements in the edge recognition result, determine the chat elements irrelevant to the chat content from the remaining chat elements in the edge recognition result to obtain at least one second candidate chat element; According to the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each second candidate chat element from the text recognition result of the first picture; Filter the matched text line recognition content and the corresponding text line coordinates from the text recognition result of the first picture.

9. The method according to claim 4, wherein The first picture is a screenshot of the instant messaging notification card, and the text recognition result includes text line coordinates and text line recognition content; Filtering the interference information irrelevant to the schedule from the text recognition result and the edge recognition result of the first picture according to the filtering rule corresponding to the picture category of the first picture includes: According to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content in each color block from the text recognition result of the first picture; For any one of the color blocks, if it is determined according to the text line recognition content in the any one of the color blocks that the any one of the color blocks does not include content related to the schedule, filter the recognition data corresponding to the any one of the color blocks from the edge recognition result and the text recognition result of the first picture.

10. The method according to any one of claims 1-9, characterized in that, The text recognition result includes text line coordinates and text line recognition content; Creating the schedule information of the first picture based on the filtered first recognition result and the filtered second recognition result includes: Based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each object from the filtered text recognition result of the first picture, where the object is a chat element or a color block; Determine the scene corresponding to the first picture according to the text line recognition content corresponding to each object and the picture category of the first picture; Based on the scene corresponding to the first picture, perform text splicing on the text line recognition content corresponding to each object to obtain spliced text; Construct a prompt according to the spliced text and the scene corresponding to the first picture, where the prompt includes scene description information for describing the scene corresponding to the first picture; Input the prompt into a natural language recognition model for processing to extract the schedule information in the first picture; Create the schedule information.

11. The method according to claim 10, wherein After displaying the schedule information, it further includes: In response to an editing operation on the schedule title of the schedule information, display a target interface that includes the spliced text; In response to a selection operation on the text line content displayed in the target interface, input the text selected by the selection operation into a title input box; In response to an editing end operation, modify the schedule title to the content input in the title input box.

12. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method according to any one of claims 1-11.

13. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the method according to any one of claims 1-11.

Citation Information

Cited By

  • Multi-dimensional session information extraction method for WeChat chat screenshot

    CN121033878A