Method for displaying schedule information, electronic equipment and readable storage medium
Through pre-trained natural language models and image processing technology, the expressive schedule information is automatically extracted and merged, solving the problem of low efficiency in manual creation of schedule information in electronic devices, and achieving fast and accurate schedule information display and editing.
Patent Information
- Application Number
- CN202311813534.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-12-25
AI Technical Summary
In the prior art, electronic devices are inefficient when quickly and accurately displaying schedule information in third-party social applications, and require users to manually create schedule information, resulting in cumbersome operations.
The pre-trained target natural language model extracts the schedule information in text information and combines the time information to represent it, reduces the amount of field attention, adopts merged representation and splitting processing, combines image recognition and edge recognition to filter out interference information, and automatically creates schedule information.
It improves the efficiency and accuracy of creating agenda information, reduces the difficulty of calculation and extraction delay, and users can quickly and accurately display and edit agenda information.
Smart Images

Figure CN120257943A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and particularly to a method for displaying schedule information, an electronic device, and a readable storage medium. Background Art
[0002] With the rapid development of terminal technology, electronic devices can usually install various types of third-party social applications, such as instant messaging applications, ticket-purchasing applications, etc. Some messages or service notifications in third-party social applications usually involve schedule information. For example, the conversation content in the chat window of an instant messaging application involves the schedule information of a certain meeting. In some scenarios, users usually need to record and display the schedule information involved in third-party social applications through electronic devices. Therefore, how to enable electronic devices to quickly and accurately display schedule information has become a research hotspot. Summary of the Invention
[0003] This application provides a method for displaying schedule information, an electronic device, and a readable storage medium, which can solve the problem of how to quickly and accurately display schedule information in related technologies. The technical solutions are as follows:
[0004] In a first aspect, a method for displaying schedule information is provided. The method includes:
[0005] In response to a schedule extraction operation, obtain target text information, where the target text information includes the schedule information to be extracted. Perform schedule extraction on the target text information through a pre-trained target natural language model to obtain first schedule information. The target natural language model can extract the schedule information in the text information and merge the time information in the text information for representation. Split the time information in the first schedule information to obtain second schedule information, where the second schedule information includes the split time information and other schedule element information in the first schedule information except the time information. Display the second schedule information.
[0006] In this way, by expressing the same amount of information with fewer fields, the number of fields that the target natural language model needs to focus on during the process of extracting schedule information is reduced, so that the computational difficulty of the target natural language model is reduced, and thus the accuracy of the output result is improved. And since the length of the schedule information to be output is reduced, the extraction delay of the schedule information can also be reduced.
[0007] As an example of the present application, the first schedule information includes field values corresponding to at least one of multiple fields. The multiple fields include a title field, a start time field, an end time field, a location field, and a recurrence period field. The field value corresponding to the title field is the title. The field value corresponding to the start time field includes one or more of a start date, a recurrence start date, and a start time point. The field value corresponding to the end time field includes one or more of an end date, a recurrence end date, and an end time point. The field value corresponding to the recurrence period field is the schedule recurrence period.
[0008] In this way, by combining the start date, the recurrence start date, and the start time point into the start time field, and combining the end date, the recurrence end date, and the end time point into the end time field, that is, representing the time information by combining it through the two fields of the start time field and the end time field, the amount of attention of the target natural language model to the fields is reduced.
[0009] As an example of the present application, in the case where the target text information includes a start time point and duration description information but does not include an end time point, the field value corresponding to the end time field further includes the duration description information. The end time point in the first schedule information is represented by the start time point and the duration description information in the target text information. The duration description information is information used to describe the schedule duration. In this way, it is beneficial for the target natural language model to perform end time point reasoning for this scenario, improving the intelligence of time reasoning.
[0010] As an example of the present application, the field value corresponding to the end time field of the first schedule information includes the duration description information, and the duration description information is information used to describe the schedule duration. In this case, the specific implementation of splitting the time information in the first schedule information to obtain the second schedule information includes: splitting out the start date from the field value corresponding to the start time field of the first schedule information as the field value corresponding to the first date field. Splitting out the end date from the field value corresponding to the end time field of the first schedule information as the field value corresponding to the second date field. The first date field and the second date field are both newly added fields and are different. Determining the field value corresponding to the split end time field according to the start time point and the duration description information in the field value corresponding to the end time field of the first schedule information.
[0011] In this way, by splitting the time information in the first schedule information to be represented by multiple fields, it is convenient to display the schedule information from multiple dimensions, improving the user experience.
[0012] As an example of this application, when the periodic schedule information is not included in the first schedule information, the first date field is the start date field, and the second date field is the end date field; when the periodic schedule information is included in the first schedule information, the first date field is the repeated start date field, and the second date field is the repeated end date field. In this way, reasonable decomposition can be carried out according to whether the first schedule information includes periodic schedules, improving the accuracy of decomposition.
[0013] As an example of this application, the specific implementation of determining the field value corresponding to the split end time field according to the start time point and duration description information in the field value corresponding to the end time field of the first schedule information includes: determining the result of adding the start time point in the field value corresponding to the end time field and the duration described by the duration description information as the field value corresponding to the split end time field. In this way, the end time point is determined by calculation, making it more intuitive for subsequent display of schedule information.
[0014] As an example of this application, the target natural language model is obtained by training the initial natural language understanding model based on a sample training set. The sample training set includes multiple groups of sample training data. Each group of sample training data includes a text training sample and a schedule information sample corresponding to the text training sample, and the time information in the schedule information sample has been merged. In this way, the target natural language model does not need to pay attention to many fields during application, and can reduce the output length of schedule information, improving the accuracy of inference and reducing the inference latency.
[0015] As an example of this application, in response to the schedule extraction operation on the first picture, through the target recognition model, the first recognition result and the second recognition result are determined. The first recognition result includes the text recognition result of the first picture, and the second recognition result includes the edge recognition result and the picture category of the first picture. The edge recognition result includes the graphic and text element attribute information or color block attribute information of the first picture. According to the picture category of the first picture, the interference information irrelevant to the schedule in the text recognition result and the edge recognition result of the first picture is filtered out. Based on the filtered first recognition result and the filtered second recognition result, the target text information is obtained. In this way, it is not necessary for the user to manually create schedule information in the calendar application, improving the efficiency of creating schedule information.
[0016] As an example of this application, the picture categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. An instant messaging chat screenshot refers to a picture obtained by taking a screenshot of the chat interface in an instant messaging application. An instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of the service notification card in an instant messaging application. An order screenshot refers to a picture obtained by taking a screenshot of the order interface in an application. Other category pictures include other pictures other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots.
[0017] Among them, in the case where the first picture is an instant messaging chat screenshot, the edge recognition result includes graphic and text element attribute information. In the case where the first picture is one of an instant messaging notification card screenshot, an order screenshot, and other category pictures, the edge recognition result includes color block attribute information.
[0018] In this way, by classifying the picture categories into the above-mentioned multiple types, it is possible to adopt a targeted method for picture processing according to the picture category during the subsequent schedule information extraction process, thereby improving the effectiveness and accuracy of schedule information extraction.
[0019] As an example of this application, the text recognition result includes text line coordinates. Correspondingly, according to the picture category of the first picture, the specific implementation of filtering out the interference information irrelevant to the schedule in the text recognition result and the edge recognition result of the first picture may include: in the case where the first picture is an instant messaging chat screenshot, filtering out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture. Filtering out the recognition data corresponding to the text lines with a line height less than the target line height in the text recognition result and the edge recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture, and the target line height is the average line height of all text lines in the first picture. Filtering out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture. In this way, by filtering out the recognition data corresponding to the skewed lines, small characters, and chat timestamps in the first picture, the interference information irrelevant to the schedule can be effectively removed, that is, the information that may interfere with the schedule information extraction can be removed, thereby improving the accuracy of subsequent schedule extraction.
[0020] As an example of this application, the graphic and text element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result further includes the text line recognition content. Accordingly, the specific implementation of filtering out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture may include: determining, from the edge recognition result, a chat element whose category is a chat timestamp to obtain at least one first candidate chat element. According to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each first candidate chat element in the text recognition result of the first picture. When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element, and the target first candidate chat element is any one of the at least one first candidate chat elements. When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filter out the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of the first picture.
[0021] In this way, first determine the chat element whose category is the chat timestamp according to the edge recognition result, then match the corresponding text line recognition content from the text recognition result, determine whether it is a chat timestamp through the target natural language model based on the matched text line recognition content, and then determine whether it is a chat timestamp according to the position of the chat element, which can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0022] As an example of this application, the specific implementation of determining whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element may include: when it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located in the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, determine to filter out the recognition data corresponding to the target first candidate chat element. Or, when it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located on the right side of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes a time point and does not include a date, determine to filter out the recognition data corresponding to the target first candidate chat element. In this way, judging whether the first candidate chat element is a chat timestamp according to the position characteristics and text line characteristics of the chat timestamp can improve the accuracy of judgment.
[0023] As an example of this application, the text recognition result includes the text line coordinates and the text line recognition content. According to the picture category of the first picture, the specific implementation of filtering out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture may include: when the first picture is a screenshot of an instant messaging notification card or an order screenshot, according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content in each color block from the text recognition result of the first picture. For any one of the color blocks, if it is determined that the content related to the schedule is not included in any one of the color blocks according to the text line recognition content in any one of the color blocks, filter out the recognition data corresponding to any one of the color blocks from the edge recognition result and the text recognition result of the first picture. In this way, by filtering out the color blocks that do not include the schedule, it is convenient to only process the color blocks that include the content related to the schedule subsequently, which can improve the data processing efficiency and the accuracy of schedule information extraction.
[0024] As an example of this application, the text recognition result includes the text line coordinates and the text line recognition content. Based on the filtered first recognition result and the filtered second recognition result, the specific implementation of obtaining the target text information may include: based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each object from the text recognition result of the filtered first picture, and the object is a chat element or a color block. Determine the scene corresponding to the first picture according to the text line recognition content corresponding to each object and the picture category of the first picture. Based on the scene corresponding to the first picture, perform text splicing on the text line recognition content corresponding to each object to obtain the target text information. In this way, by determining the scene corresponding to the first picture and performing text splicing according to the scene corresponding to the first picture, the finally obtained target text information can be made to conform to the natural language rules as much as possible.
[0025] As an example of this application, after displaying the schedule information, in response to an editing operation on the title of the schedule information, display a target interface, and the target interface includes spliced text. In response to a selection operation on the text line content displayed in the target interface, input the text selected by the selection operation into the title input box. In response to the end-of-editing operation, modify the title to the content input in the title input box. In this way, by displaying the target interface, the user can quickly modify the title by smearing, improving the user experience.
[0026] As an example of the present application, the specific implementation of obtaining the first schedule information by performing schedule extraction on the target text information through a pre-trained target natural language model may include: constructing a prompt according to the target text information and the scene corresponding to the first picture, where the prompt includes scene description information for describing the scene corresponding to the first picture, and the scene corresponding to the first picture is determined based on the first recognition result and the second recognition result. Input the prompt into the target natural language model for processing to obtain the first schedule information. In this way, by adding the scene description information of the first picture to the prompt, the target word language model can more accurately infer the schedule information.
[0027] In a second aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for displaying schedule information as described in the first aspect above is implemented.
[0028] In a third aspect, a computer-readable storage medium is provided, in which instructions are stored. When it runs on a computer, the computer is caused to execute the method for displaying schedule information as described in the first aspect above.
[0029] In a fourth aspect, a computer program product containing instructions is provided. When it runs on a computer, the computer is caused to execute the method for displaying schedule information as described in the first aspect above.
[0030] The technical effects obtained in the second, third, and fourth aspects above are similar to those obtained by the corresponding technical means in the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic framework diagram of displaying schedule information shown according to an exemplary embodiment;
[0032] Figure 2 is a schematic diagram of a schedule information sample corresponding to a text training sample shown according to an exemplary embodiment;
[0033] Figure 3 is a schematic diagram of an application scenario shown according to an exemplary embodiment;
[0034] Figure 4 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0035] Figure 5 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0036] Figure 6It is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0037] Figure 7 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0038] Figure 8 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0039] Figure 9 It is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0040] Figure 10 It is a schematic diagram of a framework for displaying schedule information shown according to another exemplary embodiment;
[0041] Figure 11 It is a schematic diagram of a framework for displaying schedule information shown according to another exemplary embodiment;
[0042] Figure 12 It is a schematic diagram of a schedule information sample corresponding to a text training sample shown according to another exemplary embodiment;
[0043] Figure 13 It is a schematic diagram of a schedule information sample corresponding to a text training sample shown according to another exemplary embodiment;
[0044] Figure 14 It is a schematic diagram of the splitting of time information shown according to an exemplary embodiment;
[0045] Figure 15 It is a schematic diagram of the splitting of time information shown according to another exemplary embodiment;
[0046] Figure 16 It is a schematic diagram of a software system of an electronic device shown according to an exemplary embodiment;
[0047] Figure 17 It is a schematic diagram of an implementation framework for creating schedule information shown according to another exemplary embodiment;
[0048] Figure 18 It is a schematic flowchart of a method for displaying schedule information shown according to an exemplary embodiment;
[0049] Figure 19 It is a schematic diagram of the processing of an instant messaging chat screenshot shown according to an exemplary embodiment;
[0050] Figure 20 It is a schematic diagram of the processing of an instant messaging notification card screenshot shown according to an exemplary embodiment;
[0051] Figure 21 It is a schematic diagram of the architecture of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the implementation manners of the present application in detail with reference to the accompanying drawings.
[0053] It should be understood that the "multiple" mentioned in the present application refers to two or more. In the description of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; the "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily mean different.
[0054] The reference to "an embodiment" or "some embodiments" etc. described in the specification of the present application means that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, the statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0055] Before introducing the method for displaying schedule information provided by the embodiments of the present application, the terms or nouns related to the embodiments of the present application will be briefly described.
[0056] One-stop office software: It can provide multiple office functions, such as document processing, spreadsheets, presentation production, project management, calendars, and emails. Users can complete multiple office tasks in the same software, improving the work efficiency of users. One-stop office software usually adopts a unified user interface design, making the switching and use between various function modules more convenient and consistent.
[0057] KV format: It refers to presenting information in a tabular format (with or without a table). For example, taking the presented information including two columns as an example, the left side is the theme and the right side is the specific content corresponding to the theme.
[0058] Global collection: It means that the user can slide up and down on the screen with three fingers to trigger the electronic device to collect information.
[0059] Magic Text: It is a function for quickly extracting text from pictures. Usually, the user can turn on or off the Magic Text switch through the path of "Settings > Smart Assistant > Magic Text", so as to turn on or off this function.
[0060] Prompt: In the embodiments of the present application, it refers to the model input data formed by adding a piece of text or instruction to the input text information when using a machine learning model, that is, it includes user prompt and system prompt. The user prompt is the input text information, and the system prompt is a piece of text or instruction added. In this way, it can guide the machine learning model to generate more accurate and targeted outputs. The system Prompt can be a question, a description, a formatted input, or even some keywords. By reasonably designing the prompt, it can guide the machine learning model to understand the intention of the question, so as to generate a relatively accurate answer. In short, Prompt is a technology widely used in machine learning models, which can help users solve data deviation, improve the controllability of machine learning models and the problem modeling ability.
[0061] Vertical domain: It refers to providing specific services for a defined group, including industries such as entertainment, medical care, environmental protection, education, and sports.
[0062] Currently, various third-party social applications can be installed in electronic devices. Exemplarily, it includes but is not limited to instant messaging (IM) applications, ticketing applications, etc. For example, the instant messaging application can be WeChat, QQ, DingTalk, Feishu, etc., and the ticketing application can be Ctrip, Qunar, Tongcheng, Fliggy, etc. The user can chat and book tickets through the third-party social application, and can also follow the official accounts (or notification accounts) of various industries through the third-party social application, so as to understand the fields and events to be concerned through the service notification cards published by the official accounts (or notification accounts). In some scenarios, some messages in the third-party social application may involve schedule information. For example, one or more messages in the chat window of the instant messaging application involve the relevant schedule of an event. The user usually has the need to record schedule information. In this case, generally, the user needs to manually create schedule information in the calendar application. However, manually creating schedule information is rather cumbersome, resulting in a low creation efficiency of schedule information.
[0063] In some embodiments, to improve the creation efficiency of schedule information, refer to Figure 1 . The electronic device obtains text information including schedule information, pre-processes the text information, and then can input the processed text information into a pre-trained model for inference to extract the schedule information, and then performs post-processing such as format conversion on the extracted schedule information, and displays the post-processed schedule information. The schedule information that users are concerned about usually may involve many fields. For example, as Figure 1 shown, the common fields involved in the schedule information include but are not limited to a title field, a start date field, an end date field, a start time field, an end time field, a location field, a repeat period field, a repeat start date field, and a repeat end date field. Among them, the field value corresponding to the title field is the title, the field value corresponding to the start date field is the start date, the field value corresponding to the end date field is the end date, the field value corresponding to the start time field is the start time point, the field value corresponding to the end time field is the end time point, the field value corresponding to the location field is the location, the field value corresponding to the repeat period field is the schedule repeat period, the field value corresponding to the repeat start date field is the repeat start date, and the field value corresponding to the repeat end date field is the repeat end date. Therefore, in the model training stage, it is necessary to construct a schedule information sample corresponding to the text training sample according to the field formats of the above multiple fields to obtain sample training data. For example, refer to Figure 2 Figure (a) in. The content of a certain text training sample is "Have a meeting at 501 from 8:00 to about two hours every Tuesday from October 1st to October 20th". In this case, when constructing the schedule information sample corresponding to this text training sample, set the field value corresponding to the title field to "Have a meeting", the field value corresponding to the start date field to be empty, the field value corresponding to the end date field to be empty, the field value corresponding to the start time field to 8:00, the field value corresponding to the end time field to 10:00, the field value corresponding to the location field to 501, the field value corresponding to the repeat period field to every Tuesday, the field value corresponding to the repeat start date field to October 1st, and the field value corresponding to the repeat end date field to October 20th. Another example, refer to Figure 2In Figure (b), the content of a certain text training sample is "Go on vacation in Xiamen from 10:18:00 to 10:20 on October 1st to October 20th". In this case, when constructing the schedule information sample corresponding to this text training sample, set the field value corresponding to the title field to vacation, the field value corresponding to the start date field to October 1st, the field value corresponding to the end date field to October 20th, the field value corresponding to the start time field to 8:00, the field value corresponding to the end time field to 10:00, the field value corresponding to the location field to Xiamen, the field value corresponding to the repeat period field to empty, the field value corresponding to the repeat start date field to empty, and the field value corresponding to the repeat end date field to empty. In this way, a large number of text training samples are obtained, and the schedule information sample corresponding to each text training sample is constructed. Each text training sample and its corresponding schedule information sample are used as a set of sample training data, and multiple sets of sample training data can be obtained. After training the initial model (such as a natural language understanding model) with this multiple sets of sample training data, a trained model can be obtained. In this way, when extracting the schedule through the trained model, the fields that the model needs to pay attention to include the title field, start date field, end date field, start time field, end time field, location field, repeat period field, repeat start date field, and repeat end date field. It can be seen that the model needs to pay attention to more fields. In this way, the computational difficulty of the model is increased, and the probability of the output result of the model being incorrect is likely to increase. Moreover, since there are more fields to be concerned about, the model needs to output the schedule information in the field formats of these multiple fields. And the model extracts the schedule information iteratively, that is, each time a token (such as a character or a word) is output, so the length of the output schedule information is relatively long, resulting in an increase in the extraction duration of the schedule information. In addition, the time information in some scenarios also needs to be inferred, and it is usually difficult for a smaller model to achieve this.
[0064] For this reason, the embodiments of the present application provide a method for displaying schedule information, which can enable an electronic device to automatically create schedule information and improve the creation efficiency of schedule information. Moreover, the embodiments of the present application increase the information density and use fewer fields to express the same amount of information by combining the Chain-of-Thought (COT) method, thereby reducing the extraction duration of schedule information, reducing the computational difficulty of the model, and improving the accuracy of the output result of the model to ensure that the field values corresponding to all fields of the schedule information are correct as much as possible.
[0065] As an example of the present application, referring to Table 1, the electronic device can create the schedules involved in the following scenarios into schedule information:
[0066] Table 1
[0067]
[0068] The above graphics and text include pictures and text information. That is, the method provided in the embodiments of the present application can not only automatically extract schedule information from text information, but also automatically extract schedule information from pictures. Specifically, the scope supported by the method provided in the embodiments of the present application is shown in Table 2:
[0069] Table 2
[0070]
[0071] As can be seen from Table 2, the electronic device can extract schedule information from pictures. The pictures can be screenshot pictures or pictures taken by a camera. The form of the screenshot picture can include but is not limited to full-screen screenshots or area screenshots. The windows involved in the screenshot picture can be full-screen, split-screen or floating windows. The sources of the screenshot pictures usually come from mobile phones, tablets, PCs, etc.; the form of the pictures taken by the camera can be but is not limited to printed text, handwritten text and artistic words. In addition, the electronic device can also extract schedule information from text. The form of the text includes but is not limited to Text and Webview. The format of the text can be plain text, formatted text or text mixed with pictures. The length of the text can be single-paragraph, multi-paragraph, short text or long text, etc. In addition, the electronic device can also extract schedule information from voice. For example, it can receive voice through a voice assistant and then convert it into text information for processing. The embodiments of the present application will be mainly described by taking the input as a picture as an example.
[0072] For ease of understanding, the application scenarios provided in the embodiments of the present application will be introduced next.
[0073] In one example, the multi-turn conversation in the chat window of WeChat has a schedule intention. When the user wants to create relevant schedule information, the user can trigger the mobile phone to take a screenshot of the application interface where the chat window is located. For example, the user can double-click on the screenshot trigger area of the mobile phone screen to trigger the screenshot operation, so that the mobile phone takes a screenshot of the interface where the WeChat chat window is located. After that, referring to Figure 3 Figure (a) therein, the mobile phone displays the screenshot editing interface U1 and displays the instant messaging chat screenshot obtained after the screenshot in the screenshot editing interface U1, that is, displays the picture p1. The screenshot editing interface U1 includes a "Share" control. When the user wants to create schedule information in the picture p1, the user can click the "Share" control. Referring to Figure 3 Figure (b) therein, in response to the user's trigger operation on the "Share" control, the mobile phone displays the sharing floating window 10. The sharing floating window 10 includes a calendar icon 11, and the user can click the calendar icon 11. In response to the user's trigger operation on the calendar icon 11, the mobile phone starts to process the picture p1 to extract and create relevant schedule information. Exemplarily, referring to Figure 3In figure (c) during this process, the mobile phone can display a prompt message "Parsing schedule information offline" so that the user can know that the schedule is being extracted from picture p1. Refer to Figure 3 In figure (d), after the schedule information is successfully created, the mobile phone displays a schedule display interface U2. The schedule display interface U2 includes a schedule display window 12, and the created schedule information 13 is displayed in the schedule display window 12. The schedule information 13 includes content such as a title, a start time point, an end time point, a start date, a location, etc. In this way, the mobile phone achieves the purpose of automatically creating and displaying schedule information, which can avoid the need for the user to manually record and improve the efficiency of creating schedule information.
[0074] As an example of this application, the schedule display window 12 also includes multiple editing controls, and the user can also edit the schedule information 13 created by the mobile phone based on the multiple editing controls according to needs. For example, refer to Figure 4 In figure (a), when the user wants to modify the title of the schedule, they can click on the title editing control 14 in the schedule display window 12. In response to the user's click operation on the title editing control 14, the mobile phone displays a target interface U3 (which can be called a smearing interface) as shown in Figure 4 figure (b). The target interface U3 displays information related to the schedule in picture p1. In this way, the user can smear on the information displayed in the target interface U3 to modify or fill in the title of the schedule. Correspondingly, the mobile phone enters the content smeared by the user in the title input box 15 of the target interface U3. For example, refer to Figure 4 In figure (b), when the user wants to modify the title of the schedule to "HarmonyOS Talk", the user can smear the content of "HarmonyOS", "Talk", and "Open" in the target interface U3 in sequence. Correspondingly, the mobile phone enters "HarmonyOS", "Open", and "Talk" in the title input box 15 in sequence. Refer to Figure 4 In figure (c), after the user finishes smearing, they can trigger the "Input" control (or "√" control) of the target interface U3. In response to the user's triggering operation on the "Input" control (or "√" control), the mobile phone resumes displaying the schedule display window 12. At this time, the user can see from the schedule display window 12 that the title of the schedule has been modified to the content modified by the user through the smearing operation, that is, the title of the schedule has been changed from "Talk KaiTan" to "HarmonyOS Talk". In this way, by displaying information related to the schedule and that can be smeared in the target interface U3, the user can quickly modify the title of the schedule by smearing, improving the user experience.
[0075] In addition, refer to Figure 4In Figure (a), the schedule display window 12 further includes a time editing control. When the user wants to edit the time information in the schedule information 13, the user can also modify it based on the time editing control. For example, the user can click on the displayed time information to edit it. Additionally, after the user swipes down the schedule display window 12, the schedule display window 12 can also provide other editing controls, such as editing controls for the number of repetitions, reminder time, important reminder, etc. Thus, the user can edit the schedule information based on other editing controls, and the embodiments of the present application do not limit this.
[0076] As an example of the present application, after the user clicks on the "√" control in the schedule display window 12, in response to this trigger operation, the mobile phone displays the schedule information 13 in the schedule details area of the calendar application, so that the user can view the schedule information 13 from the schedule details area of the calendar application. As an optional example, after the user clicks on the "√" control in the schedule display window 12, the mobile phone can also display the schedule information 13 in the form of a card at positions such as the desktop, the negative first screen, or the notification center, for the user to quickly view later. The embodiments of the present application do not limit this.
[0077] It should be noted that the above is only an example of the user triggering the mobile phone to create schedule information through the sharing entry (i.e., the sharing control). In another example, see Figure 3 Figure (a). The mobile phone provides a Magic text control 00 in the screenshot editing interface U1. When the user needs the mobile phone to create schedule information based on the picture p1, the user can also click on the Magic text control 00, thereby triggering the mobile phone to create and display the schedule information with one key.
[0078] In another example, the user can also trigger the mobile phone to create and display schedule information through the Any Door entry. Exemplarily, see Figure 5 Figure (a). After the mobile phone takes a screenshot of the interface where the WeChat chat window is located, the obtained picture p1 is automatically saved to the gallery. Thus, when the user wants the mobile phone to automatically create the schedule information in the picture p1, the user can open the screenshot picture interface U4 in the gallery, and the picture p1 is displayed in the screenshot picture interface U4. See Figure 5 Figure (b). The user can trigger the mobile phone to select the picture p1, and then, the user can drag the picture p1 to the right side of the mobile phone screen. When the user drags to a certain position, in response to the user's drag operation, the mobile phone displays the application programs that can receive and process the picture p1, such as Figure 5 shown in Figure (b), displaying the calendar, WeChat, and QQ. Thus, the user can continue to drag the picture p1 onto the calendar application and then release it. In response to the user's release operation, the mobile phone starts to process the picture p1 to extract and create relevant schedule information. See Figure 5In figure (c), after the mobile phone creates schedule information, the schedule information is displayed in the schedule display window 12.
[0079] In another example, the user can only select the content related to the schedule in picture p1, and then trigger the mobile phone to extract schedule information from the selected part. For example, see Figure 6 In figure (a), the user can select part of the content in picture p1. For example, in the screenshot editing interface, a control for triggering the selection operation can be provided. After triggering this control, the user can select on picture p1. In response to the user's selection operation, the mobile phone selects the part selected by the user. As an example, the mobile phone can take a screenshot of the area selected by the user, as shown in Figure 6 figure (b). After that, the user can trigger the mobile phone to extract schedule information from the selected area through an interaction entry such as the "Share" control, and create and display schedule information related to the content in this area. In this way, by supporting the user to select on picture p1, the data processing volume of the mobile phone can be reduced, thereby improving the creation efficiency of schedule information.
[0080] It should be noted that the interaction entries used by the user to trigger the mobile phone to extract schedule information from picture p1 in the above application scenarios are only exemplary. In some embodiments, the mobile phone can also be triggered to extract schedule information from picture p1 through other interaction entries. For example, it can also be triggered through interaction entries such as global collection. The embodiments of the present application do not limit this.
[0081] The above application scenarios are only exemplary. In addition, the mobile phone can also extract schedule information from some types of pictures in other scenarios. Exemplarily, see Figure 7 In figure (a), there is a forwarded picture p2 (which can be obtained by screenshot or shooting) in a certain chat window, and there is schedule information in picture p2. When the user wants to extract the schedule information in picture p2, the user can drag picture p2 to the right side of the mobile phone screen. See Figure 7 In figure (b), when the user drags to a certain position, in response to the user's drag operation, the mobile phone displays application programs that can receive and process picture p2. For example, as shown in Figure 7 figure (b), the calendar, WeChat, and QQ are displayed. After the user drags picture p2 onto the calendar application and releases it, in response to the user's release operation, the mobile phone starts to automatically process picture p2 to extract and create relevant schedule information. See Figure 7 In figure (c), after the mobile phone creates schedule information, the schedule information is displayed in the schedule display window 12.
[0082] See Figure 7In figure (c), when the mobile phone creates schedule information multiple times, multiple schedule information can be displayed in the schedule display window 12. When not all schedule information can be fully displayed in the schedule display window 12, the user can trigger the mobile phone to display the hidden schedule information by swiping the schedule display window 12 left and right.
[0083] In addition, the mobile phone not only supports the user to drag the pictures in the chat window to the entrance of the Any Door, but also supports the user to drag the text to the entrance of the Any Door. For example, a certain chat window includes chat text, and the chat text includes schedule information. When the user needs the mobile phone to create the schedule information in the chat text, the user can select the chat text, and then Figure 7 drag the chat text to the calendar application in the manner shown. Correspondingly, the mobile phone extracts the schedule information in the chat text, creates and displays the relevant schedule information.
[0084] In another example, the mobile phone can also extract schedule information from the screenshot of the instant messaging notification card. For example, refer to Figure 8 figure (a). The figure shows a schematic diagram of a screenshot of an instant messaging notification card (i.e., picture p3) shown according to an exemplary embodiment. The picture p3 is obtained by taking a screenshot of the notification card published in the service notification in WeChat, and it includes schedule information. When the user wants the mobile phone to create and display the schedule information in the picture p3, the user can trigger the mobile phone to extract the schedule information according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information from the picture p3 through the sharing entrance. Correspondingly, the mobile phone extracts the schedule information based on the picture p3, and then creates or displays the relevant schedule information. For example, refer to Figure 8 figure (b). The mobile phone creates and displays schedule information 60, and the schedule information 60 includes the theme, ticket number, departure time point, departure date, location, etc.
[0085] In another example, the mobile phone can also extract schedule information from the order screenshot, and the order screenshot can be a screenshot picture such as a hotel order, a train ticket order, an airplane ticket order, etc. For example, refer to Figure 9 figure (a). The figure shows a schematic diagram of an order screenshot (i.e., picture p4) shown according to an exemplary embodiment. The picture p4 is obtained by taking a screenshot of the interface where the train ticket order is located, and the picture p4 includes schedule information. When the user wants the mobile phone to create and display the schedule information in the picture p4, the user can trigger the mobile phone according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information from the picture p4 through the Magic text entrance. Correspondingly, the mobile phone extracts the schedule information based on the picture p4, and then creates or displays the relevant schedule information. For example, the displayed schedule information is as shown in Figure 9 70 in figure (b). The schedule information 70 includes the theme, train number, departure time point, arrival time point, departure date, location, etc.
[0086] It should be noted that the above application scenarios are all exemplary and do not limit the application scenarios of the method provided by the embodiments of the present application. In another embodiment, the mobile phone can also extract schedule information from other types of pictures, and other types of pictures include pictures other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots. Of course, other types of pictures can be pictures obtained by taking screenshots or pictures taken by a camera.
[0087] Please refer to Figure 10 , Figure 10 which is a schematic diagram of the implementation framework of a method for displaying schedule information shown according to an exemplary embodiment. In implementation, the electronic device obtains text information, and the text information can be from pictures, text, or voice. In the case of being from pictures, the electronic device can obtain text information by performing image recognition processing on the pictures, and the user operation process can refer to the above; in the case of being from text, the electronic device can obtain text information through operations such as copying; in the case of being from voice, the electronic device can obtain text information by performing speech recognition processing on the voice. After that, the electronic device can extract schedule information through a pre-trained target natural language model, and after post-processing the extracted schedule information, display the processed schedule information.
[0088] As an example of the present application, the target natural language model can extract the schedule information in the text information and merge and represent the time information in the text information, that is, the time information in the schedule information extracted by the target natural language model is merged and represented, rather than being represented separately in multiple fields. For example, refer to Figure 11 , the schedule information output by the target natural language model includes the field values corresponding to the title field, start time field, end time field, location field, and repeat period field respectively, that is, the time information in the text information is merged into two fields (start time field and end time field) for encoding representation. In comparison, the number of fields involved in the schedule information output by the target natural language model is reduced.
[0089] As an example of the present application, the target natural language model is pre-trained. In one example, the target natural language model is obtained by training an initial natural language understanding model based on a sample training set. The sample training set includes multiple groups of sample training data, and each group of sample training data includes a text training sample and a schedule information sample corresponding to the text training sample, and the time information in the schedule information sample is processed by merging.
[0090] As an example, the initial natural language model is a large language model (LLM). Exemplarily, the LLM can be a natural language processing (NLP) model.
[0091] During the training process, when setting the schedule information sample corresponding to each text training sample, the time information can be merged to obtain the processed schedule information sample. As an example of the present application, the start date field, start time field, and repeated start time field can be merged, and the end date field, end time field, and repeated end time field can be merged. Additionally, if the text training sample does not include an end time point but includes a start time point and duration description information, where the duration description information is used to describe the schedule duration, then when constructing the schedule information sample, the field value in the end time field can be set to the start time point and duration description information in the text training sample.
[0092] Exemplarily, referring to Figure 12 , for a text training sample including periodic schedule information such as "Have a meeting at 8:00 every Tuesday for about two hours from October 1st to October 20th in Room 501", its corresponding schedule information sample can be modified from the original structure shown in Figure (a) of Figure 12 to the structure shown in Figure (b) of Figure 12 , that is, merging the start date (empty) corresponding to the start date field and the repeated start date (October 1st) corresponding to the repeated start date field into the field value corresponding to the start time field, and merging the end date (empty) corresponding to the end date field and the repeated end date (October 20th) corresponding to the repeated end date field into the field value corresponding to the end time field. The end time point in the end time field is represented by "8:00" and "about two hours". That is, when the text training sample includes a start time point and duration description information but does not include an end time point, when setting its corresponding schedule information sample, instead of calculating the end time point based on the start time point and duration description information in the text training sample, the end time point is represented by the start time point and duration description information in the text training sample for the model to learn, so that in subsequent model applications for the same scenario, the target natural language model can automatically infer the end time point, improving the intelligence of the model.
[0093] Another example, referring to Figure 13, for a text training sample such as "10.1 8:00 - 10.20 10:00 go on vacation in Xiamen" that does not include periodic schedule information, the corresponding schedule information sample can be modified from the structure shown in Figure (a) of Figure 13 to the structure shown in Figure (b) of Figure 13 , that is, merge the start date corresponding to the start date field (October 1st) and the repeated start date corresponding to the repeated start date field (empty) into the field value corresponding to the start time field, and merge the end date corresponding to the end date field (October 20th) and the repeated end date corresponding to the repeated end date field (empty) into the field value corresponding to the end time field. Similarly, if the text training sample includes start time point and duration description information but does not include the end time point, when setting the corresponding schedule information sample, the end time point is not calculated based on the start time point and duration description information in the text training sample, but the end time point is directly represented by the start time point and duration description information in the text training sample.
[0094] When merging and representing time information, time points and dates can be divided by a preset identifier. For example, the preset identifier can be "|", etc., and the date and time point can be arranged according to certain rules. For example, the date is arranged before the time point. The embodiments of the present application do not limit this.
[0095] The number of preset identifiers can be set according to requirements. For example, in the case of duration description information, two preset identifiers can be included, as shown in Figure (b) of Figure 12 . Or only 1 can be included. In this case, the time information before the preset identifier can be the date and time point, and the time information after the preset identifier is the duration description information. In the case of no duration description information, it can be divided by 1 identifier.
[0096] As an example of the present application, if the text training sample includes the end time point and duration description information but does not include the start time point, when constructing the schedule information sample corresponding to the text training sample, the start time point can be represented by the end time point and duration description information in the field value corresponding to the start time field. If the text training sample includes the start date and duration description information but does not include the end date, when constructing the schedule information sample corresponding to the text training sample, the end date can be represented by the start date and duration description information in the field value corresponding to the end time field. If the text training sample includes the end date and duration description information, when constructing the schedule information sample corresponding to the text training sample, the start date can be represented by the end date and duration description information in the field value corresponding to the start time field.
[0097] As an example of the present application, if the text training sample includes an end time point and a time interval duration, where the time interval duration refers to the interval duration from the end time point to the next start time point, then when constructing the schedule information sample corresponding to the text training sample, the next start time point can be represented in the field value corresponding to the start time field by the end time point and the time interval duration. For example, if the text training sample is "This meeting ends at 10:00 and resumes two hours later", then the field value corresponding to the start time field in its corresponding schedule information sample can be: |10:00| two hours later.
[0098] As an example of the present application, if the text training sample does not include a start time point and an end time point, but only includes a time interval duration, then when constructing the schedule information sample corresponding to the text training sample, only the time interval duration can be included in the field value corresponding to the start time field. For example, if the text training sample is "Starts two hours later", then the field value corresponding to the start time field in its corresponding schedule information sample can be: || two hours later.
[0099] Set the schedule information sample corresponding to each text training sample in this way. Then, take each text training sample and its corresponding schedule information sample as a set of sample training data to obtain multiple sets of sample training data. After iteratively training the initial natural language model with the multiple sets of sample training data, the target natural language model can be obtained. In this way, during the application process, after inputting the text information into the target natural language model, the target natural language model will output the schedule information according to the modified field specifications, for example, output according to the field formats corresponding to the title field, start time field, end time field, location field, and repeat cycle field.
[0100] It is worth mentioning that expressing the same amount of information with fewer fields increases the information density, reducing the number of fields that the target natural language model needs to pay attention to during the process of extracting schedule information. As a result, the computational difficulty of the target natural language model is reduced, thereby improving the accuracy of the output result. And since the length of the schedule information to be output is reduced, the extraction latency of the schedule information can also be reduced.
[0101] Since the time information in the schedule information output by the target natural language model during the application process is represented in a combined manner, that is, using fewer fields, therefore, refer to Figure 11, during the post - processing, the electronic device decodes the structure of the schedule information output by the target natural language model to split the combined time information, so as to represent it one by one using more fields. As an example of this application, during the splitting process, the start date is split from the field value corresponding to the start time field of the schedule information output by the target natural language model as the field value corresponding to the first date field. The end date is split from the field value corresponding to the end time field of the schedule information output by the target natural language model as the field value corresponding to the second date field. The first date field and the second date field are both newly added fields and are different. In addition, when the field value corresponding to the end time field of the schedule information output by the target natural language model includes duration description information, the electronic device can also determine the field value corresponding to the split end time field according to the start time point and the duration description information in the field value corresponding to the end time field of the schedule information output by the target natural language model.
[0102] As an example but not a limitation, the time information can be split according to a preset identifier. For example, for the field value corresponding to the start time field, the time information before the first preset identifier is the start date, the time information between the first preset identifier and the second preset identifier is the start time point, and the time information after the second preset identifier is the duration description information.
[0103] When the schedule information output by the target natural language model does not include periodic schedule information, the first date field is the start date field, and the second date field is the end date field. The field value corresponding to the start date field is the date when the schedule starts, and the field value corresponding to the end date field is the date when the schedule ends; when the schedule information output by the target natural language model includes periodic schedule information, the first date field is the repeated start date field, and the second date field is the repeated end date field. The field value corresponding to the repeated start date field is the date when the schedule repeats start, and the field value corresponding to the repeated end date field is the date when the schedule repeats end. As an example but not a limitation, the electronic device can determine whether the schedule information includes periodic schedule information according to whether the field value corresponding to the repeat period field is empty.
[0104] As an example of this application, according to the start time point and duration description information in the field value corresponding to the end time field of the schedule information output by the target natural language model, determine the field value corresponding to the split end time field, including: determining the result of adding the start time point in the field value corresponding to the end time field to the duration described by the duration description information as the field value corresponding to the split end time field. Specifically, the electronic device can identify time entities. For example, it can be identified by a rule matching method. For example, when the time information satisfies xx:xx, it is determined as a time point, and the other is the duration description information, so as to determine the start time point and the schedule duration, and then add the start time point to the schedule duration to obtain the end time point.
[0105] Of course, in another example, the electronic device can also input the field value corresponding to the split end time field into other models for separate time reasoning. For example, input "8:00 about two hours" into other models for time reasoning to determine the end time point. The other model can be the target natural language model or other models, and the embodiments of this application do not limit this.
[0106] Exemplarily, refer to Figure 14 and the structure of the schedule information output by the target natural language model is as shown in Figure 14 Figure (a) therein, which includes the field value corresponding to the title field (for a meeting), the field value corresponding to the start time field (for October 1st | 8:00), the field value corresponding to the end time field (for October 20th | 8:00 | about two hours), the field value corresponding to the location field (for 501), and the field value corresponding to the repeat period field (for every Tuesday). In this case, during the splitting process, the start date (October 1st) in the field value corresponding to the start time field is used as the field value corresponding to the repeat start date field, the end date (October 20th) in the end time field is used as the field value corresponding to the repeat end date field, and 8:00 in the field value corresponding to the end time field is added to 2 hours to determine the field value finally corresponding to the end time field. In this way, the structure shown in Figure 14 Figure (b) therein is obtained, which is the split schedule information.
[0107] Exemplarily, refer to Figure 15 and the structure of the schedule information output by the target natural language model is as shown in Figure 15As shown in Figure (a), it includes the field value corresponding to the title field (for playing basketball), the field value corresponding to the start time field (for Monday), the field value corresponding to the end time field (empty), the field value corresponding to the location field (empty), and the field value corresponding to the repeat period field (for every Monday). That is, the schedule information output by the target natural language model includes the field values corresponding to at least one of the above multiple fields. In this case, during the splitting process, the start date in the field value corresponding to the start time field is used as the field value corresponding to the repeat start time field. In this way, a structure as shown in Figure (b) of Figure 15 can be obtained, which is the split schedule information.
[0108] As an example of this application, if the field value corresponding to the start time field of the schedule information output by the target natural language model includes an end time point and a time interval duration, the electronic device determines the start time point based on the end time point and the time interval duration. For example, if the field value corresponding to the start time field of the schedule information output by the target natural language model is: |10:00| two hours later, the start time point can be determined to be 12:00.
[0109] As an example of this application, if the field value corresponding to the start time field of the schedule information output by the target natural language model only includes a time interval duration, the electronic device determines the start time point based on the current system time and the time interval duration. For example, if the current system time is 5:00 and the field value corresponding to the start time field of the schedule information output by the target natural language model is: || two hours later, the start time point can be determined to be 7:00.
[0110] In this way, the schedule information output by the target natural language model can be restored to the field format required in daily life. After that, the schedule information can be displayed based on the split structure.
[0111] It should be noted that the above implementation method of outputting schedule information through the target natural language model and then splitting the schedule information output by the electronic device is only exemplary. In another example, when training the target natural language model, schedule information samples corresponding to the split field format can also be added to the sample training data. That is, the sample training data not only includes text training samples and merged schedule information samples, but also includes split schedule information samples. For example, referring to Figure 12 , the sample training data not only includes text training samples and schedule information samples as shown in Figure (b) of Figure 12 , but also includes those as shown in Figure 12The sample schedule information shown in Figure (a) therein. In this way, during the model application process, the target natural language model can not only output the merged schedule information, but also, after inputting the merged schedule information into the target natural language model, the target natural language model can further split the merged schedule information to output the split schedule information. For example, after inputting the schedule information shown in Figure (a) as Figure 14 therein into the target natural language model, it can output the schedule information shown in Figure (b) as Figure 14 therein. In this way, the COT method is adopted to extract the schedule information.
[0112] Of course, it should be noted that the embodiments of this application illustrate by taking the extraction of schedule information through the target natural language model as an example. In another example, if the text information satisfies certain rules, the schedule information can also be extracted by the rule matching method. For example, if the text information satisfies the rule of removing "xx" from "xx time", the information before "go" can be extracted as time information, and the information after "go" can be extracted as the location. In this case, the electronic device can also extract the schedule information according to the above field rules, and the embodiments of this application do not limit this.
[0113] The software system of the electronic device involved in the embodiments of this application can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. The embodiments of this application take the Android system with a layered architecture as an example to exemplarily illustrate the software system of the electronic device.
[0114] Figure 16 is a block diagram of the software system of an electronic device provided by the embodiments of this application. Refer to Figure 16 , the layered architecture divides the software into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, which are the application layer, the application framework layer, the Android runtime and the system layer, and the kernel layer.
[0115] The application layer may include a series of application packages. Refer to Figure 16 , the application packages may include a calendar and other applications. For example, other applications may include instant messaging, ticket booking, camera, gallery, call, map, navigation, Bluetooth, music, video, short message, voice assistant and other applications.
[0116] In addition, as an example of the present application, the application layer further includes a schedule management service and a model management service. The schedule management service can be used to provide services for the calendar, or to be called by the calendar. For example, the schedule management service can create schedule information for the calendar and store data such as schedule information. The schedule management service can include a schedule creation service (which can be called: intelligent parsing and processing service) and a schedule database (such as: Calednar Provider schedule database). The schedule management service can create schedule information through the schedule creation service and store the created schedule information through the schedule database. The model management service (which can be called: MagicLive large model service) can be used to provide various models for the schedule management service to call when needed.
[0117] In one example, the model management service provides a pre-trained large language model (LLM) model, an object recognition model, and a personal behavior feature model. Among them, the LLM model can be used to identify (i.e., infer) the prompt to determine schedule information. Exemplarily, the LLM model can be a natural language processing (NLP) model. The NLP model can run through a natural language unit (NLU). In the embodiments of the present application, the pre-trained LLM model is referred to as the target natural language model. In some examples, the target natural language model can not only identify the prompt but also extract keywords in a text, such as keywords like time and location. The object recognition model can be used to perform text recognition and edge recognition on pictures, and can also be used to determine the picture category of the pictures. In one example, the object recognition model includes a first optical character recognition (OCR) model, a second OCR model, and an edge detection model. The first OCR model can be used for text recognition, the second OCR model can be used to determine the picture category, and the edge detection model can be used to perform edge recognition on pictures. The number of edge detection models can be multiple, and different edge detection models can be used to perform edge recognition on pictures of different picture categories; the personal behavior feature model can be used to determine a user profile based on the user's historical behavior data.
[0118] It should be noted that the embodiments of the present application are described by taking the target natural language model as a pre-trained LLM model as an example. In another example, the size of the model and whether it is multimodal may not be limited.
[0119] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, Figure 16 As shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.
[0120] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system. The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0121] The system libraries may include multiple functional modules, such as: surface manager, Media Libraries, 3D graphics processing libraries (such as: OpenGL ES), 2D graphics engines (such as: SGL), etc.
[0122] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0123] The electronic device can implement the method for displaying schedule information provided in the embodiments of the present application through the interaction of the above multiple modules. Exemplarily, taking the object to be extracted with schedule information as a picture as an example, see Figure 17, in implementation, the electronic device can transfer the picture to be processed to the calendar application through interaction entrances such as sharing, portal, global collection, or Magic text. Among them, the picture can be a screenshot of the order of train tickets, air tickets, or a screenshot of the order of hotels, catering, entertainment, or a screenshot of the order of sports, health, or a screenshot of a WeChat mini-program, web, etc., or a screenshot of the service notification card of the IM notification number, or a screenshot of the IM chat conversation, etc. By way of example and not limitation, the screenshot of the order of train tickets or air tickets can be from applications such as 12306, China Railway, or Ctrip, and the screenshot of the order of hotels, catering, entertainment can be from applications such as Ctrip, Qunar, Tongcheng, Meituan, Fliggy, Dianping, Damai, etc., and the screenshot of the order of sports, health can be from applications such as Keep, Lianduoduo, registration platform, etc., and the screenshot of a WeChat mini-program, web, etc. can include content such as performances, flash sales, marathons, etc., and the screenshot of the service notification card of the IM notification number can include notification cards issued by third parties such as hospitals, scenic spot tickets, educational institutions, insurance companies, etc. through the IM notification number, and the screenshot of the IM chat conversation can include content such as work arrangements, invitations, educational tasks, etc.
[0124] After the calendar application receives the picture, it requests the target recognition model to perform text recognition and edge recognition on the picture. Through text recognition, the text information in the picture can be extracted, and through edge recognition, the chat elements or color blocks in the picture can be recognized. For example, see Figure 17 , transfer the picture p1 to the target recognition model through the interaction entrance. After text recognition by the target recognition model, the text recognition result is as shown in Figure 17 90. In addition, after edge recognition by the target recognition model, the coordinates and categories of each chat element in the picture p1 can be recognized.
[0125] After that, the electronic device can perform filtering processing on the text recognition result and the edge recognition result based on the filtering rules, that is, perform pre-processing to remove the interference information unrelated to the schedule in the text recognition result and the edge recognition result. As an example of the application, the preset filtering rules can include at least one of the following rules: 1. Discard dense text blocks. 2. Discard small characters. 3. Discard skewed lines. 4. Discard floating text on the attached drawing. 5. Discard the "back" in the upper left corner of the picture and the "+" symbol in the picture. Exemplarily, after filtering the text recognition result based on the filtering rules, the content shown in Figure 17 91 can be removed. In addition, after filtering the edge recognition result based on the filtering rules, the coordinates and categories of the chat elements unrelated to the schedule can be removed.
[0126] In the process of extracting the schedule from the picture, if the text recognition result is not filtered, but directly input into the target natural language model for schedule information extraction. In this case, referring to Table 3, the recognition success rate of the target natural language model is relatively low, usually only reaching about 30%. The reason is the problems of interference information, format line breaks, and loss of layout information. Therefore, in the embodiments of the present application, filtering the text recognition result before extracting the schedule information through the target natural language model can improve the accuracy of the subsequent target natural language model in extracting schedule information.
[0127] Table 3
[0128]
[0129] After the filtering process, the electronic device splices the target text information based on the remaining text recognition result and the edge recognition result after filtering, constructs a prompt (i.e., prompt) corresponding to the target text information. Then, the prompt is input into the target natural language model to extract schedule information through the target natural language model. The target natural language model outputs schedule information based on the input prompt. For example Figure 17 92 in shows a part of the content processed by the target natural language model. Then the electronic device splits the time information of the output schedule information, and performs post-processing on the schedule fields of the split schedule information to obtain the schedule information to be displayed. Exemplarily, the schedule information to be displayed is as shown in Figure 17 93 in. After that, the electronic device can display the schedule information.
[0130] In one example, the post-processing of the schedule field may include but is not limited to at least one of the following: 1. Processing according to the reminder time rule. 2. Processing according to the start time rule of tomorrow. 3. Processing according to the details rule. 4. Processing based on the pre-filled rule of title merging.
[0131] The reminder time rule includes setting the reminder time of the schedule information earlier than the time information extracted by the target natural language model by a preset duration. The preset duration can be set according to requirements. For example, the preset duration is 30 minutes. If the start time point extracted by the target natural language model is 8:30, the reminder time of the schedule information can be set to 8:00. In addition, the reminder time rule also includes repeated reminders. For example, if the schedule information extracted by the target natural language model is to grab numbers on Monday, Tuesday, and Wednesday, the electronic device sets the reminder time for Monday, the reminder time for Tuesday, and the reminder time for Wednesday for this schedule information, rather than setting only one reminder time.
[0132] The tomorrow start time rule means that if there is a situation of crossing days, months or years in the time information extracted by the target natural language model, then the time after crossing days, months or years is supplemented. For example, if the date extracted by the target natural language model is December 5th, the start time point is 23:00, and the end time point is 00:30, then the electronic device can supplement the end time point as 00:30 on December 6th in the schedule information.
[0133] The details rule means that the layout and font size of the created schedule information are adjusted according to the size of the screen of the electronic device, so that it can be correctly and clearly displayed in the schedule details area of the calendar application.
[0134] The title merging pre-filling rule includes, in the case where there are multiple different titles corresponding to the same time information, selecting the title with the longest length as the title of the schedule information from the multiple titles. In addition, the title merging pre-filling rule also includes using the specified title corresponding to the scene of the picture as the title of the schedule information. For example, if the scene of the picture is making an appointment to get a number for seeing a doctor, and the title extracted by the target natural language model is "getting a number" or "fetching a number", then the title of the schedule can be standardized as "registering for seeing a doctor", where the specified titles corresponding to different scenes can be set in advance according to requirements.
[0135] It should be noted that the above post-processing of the schedule fields is only exemplary. In another example, the post-processing of the schedule fields may also include but is not limited to at least one of time similarity processing, discarding empty results, risk control, and cleaning non-natural language titles. Time similarity processing means that if the time in the output schedule information is earlier than the current system time of the electronic device, then the time closest to the time in the schedule information is determined according to the current system time, and the determined time is determined as the time in the schedule information. For example, if the time in the schedule information is Tuesday and the current system time is Wednesday, then it is recorded as Tuesday of the next week in the schedule information. Discarding empty results means discarding the returned empty fields. Risk control means controlling sensitive words. Cleaning non-natural language titles means that if the title of the schedule does not include a verb, a verb can be added to the title of the schedule or the title of the schedule can be default set according to the scene corresponding to the picture.
[0136] Next, in combination with Figure 18 The process of extracting and displaying the schedule information in the picture will be introduced in detail. See Figure 18 This method may include the following implementation steps:
[0137] S1801: The calendar application receives the picture L to be processed.
[0138] The picture L can be a picture in bitmap format.
[0139] The calendar application receives picture L passed by the interaction entry. As described above, the interaction entry can be a sharing entry, an arbitrary door entry, etc. Exemplarily, referring to Figure 3 Figure (a) in, when the user submits picture L to the calendar application through the sharing control, this interaction entry is the sharing entry.
[0140] S1802: The calendar application sends a schedule creation instruction to the schedule management service, and picture L is carried in the schedule creation instruction.
[0141] The schedule creation instruction is used to indicate to create and display relevant schedule information based on picture L.
[0142] S1803: The schedule management service sends picture L to the first OCR model in the model management service.
[0143] In implementation, the schedule management service calls the first OCR model in the model management service and sends picture L to the first OCR model for text recognition processing.
[0144] S1804: The first OCR model determines the first recognition result of picture L.
[0145] As an example of this application, the first recognition result includes the text recognition result of picture L, and the text recognition result includes text block coordinates, text line coordinates, and text line recognition content.
[0146] As an example, the first recognition result further includes target indication information, which can be used to indicate whether the picture L input into the first OCR model is a screenshot picture or a taken picture. Exemplarily, the target indication information can be a first identifier, a second identifier, or a third identifier. The first identifier is used to indicate that the picture L input into the first OCR model is a screenshot picture. The second identifier is used to indicate that the picture L input into the first OCR model is a taken picture and is a picture taken of a document. The third identifier is used to indicate that the picture L input into the first OCR model is other taken pictures, such as pictures taken of advertisements, road signs, magazines, etc. The first identifier, the second identifier, and the third identifier can be set according to requirements. For example, the first identifier is F1, the second identifier is F2, and the third identifier is F3.
[0147] That is, after receiving picture L, the first OCR model recognizes picture L and outputs the first recognition result of picture L.
[0148] S1805: The first OCR model sends the first recognition result to the schedule management service.
[0149] As an example, after receiving the first recognition result, the schedule management service can cache the first recognition result.
[0150] S1806: The schedule management service sends picture L to the second OCR model in the model management service.
[0151] In one example, after the schedule management service receives a schedule creation instruction sent from a calendar application, in addition to sending picture L to the first OCR model for text recognition processing, it can also call the second OCR model in the model management service and send picture L to the second OCR model to determine the picture category through the second OCR model. That is, the operations of S1806 and S1803 can be executed in parallel.
[0152] Since the method provided by the embodiments of this application can extract schedule information from pictures of different picture categories, and the content layouts of pictures of different picture categories are different, the electronic device processes pictures of different picture categories in different ways. Therefore, in implementation, after the schedule management service receives the to-be-processed picture L, it not only performs text recognition through the first OCR model, but also inputs picture L into the second OCR model to determine the picture category of picture L.
[0153] As an example of this application, the picture categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. An instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application; an instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application; an order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application; other category pictures include other pictures except instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots, and other category pictures can be screenshot pictures or captured pictures.
[0154] S1807: The second OCR model determines the picture category of picture L.
[0155] After receiving picture L, the second OCR model recognizes picture L and outputs the picture category of picture L.
[0156] S1808: The second OCR model sends the picture category of picture L to the schedule management service.
[0157] S1809: The schedule management service sends picture L to the edge detection model corresponding to the picture category of picture L.
[0158] As an example of this application, multiple edge detection models are provided in the model management service. The multiple edge detection models can be pre-trained, and different edge detection models can identify the edges of pictures of different picture categories. In one example, the multiple edge detection models include a first edge detection model and a second edge detection model. The first edge detection model can be used to identify the edges of instant messaging chat screenshots to determine the coordinates and categories of chat elements in the instant messaging chat screenshots; the second edge detection model can be used to identify the edges of pictures other than instant messaging chat screenshots. For example, the second edge detection model can be used to identify the edges of instant messaging notification card screenshots, order screenshots, or other category pictures to determine the coordinates and categories of color blocks in other pictures.
[0159] The different edge detection models can be obtained by iteratively training an initial training model based on picture training samples of pictures of the corresponding picture categories. The picture training samples can be obtained by edge annotation in advance according to requirements. The initial training model can be set according to requirements. As an example but not a limitation, the initial training model can be a recurrent neural network (RNN), etc.
[0160] After the schedule management service receives the picture category of picture L, it determines the edge detection model corresponding to the picture category of picture L from the multiple edge detection models. Exemplarily, when picture L is Figure 3 p1 in Figure (a) of Figure 8 , that is, in the case of an instant messaging chat screenshot, the edge detection model corresponding to the picture category of picture L is determined from the multiple edge detection models as the first edge detection model; when picture L is Figure 9 p3 in Figure (a) of
[0161] S1810: The edge detection model performs edge recognition processing on picture L and outputs an edge recognition result.
[0162] The edge recognition result includes graphic and text element attribute information or color block attribute information.
[0163] In one example, when the picture category of picture L is an instant messaging chat screenshot, the edge detection model corresponding to the picture category of picture L is the first edge detection model. After inputting picture L into the first edge detection model for processing, the output edge recognition result includes graphic and text element attribute information. The graphic and text element attribute information includes the attribute information of the chat elements in the instant messaging chat screenshot. The attribute information of the chat elements may include the coordinates and categories of the chat elements. Exemplarily, the chat elements include avatars, titles, nicknames, chat content, chat timestamps, usernames, specified identifiers, etc. The specified identifier includes the "+" identifier. For example, see Figure 19 , after inputting picture L into the first edge detection model, the first edge detection model can determine that the chat elements in picture L include Figure 19 the multiple items identified by the dashed boxes in
[0164] . Each chat element corresponds to its own coordinates and category. For example, the coordinates of a certain chat element are the coordinates of the four corners of the area where the chat element is located, and the category is an avatar. Optionally, the attribute information of each chat element may further include the serial number of the chat element. The serial numbers of the chat elements in picture L can be set by default according to the sorting of the chat elements in picture L. Figure 20 , when picture L is an instant messaging notification card screenshot, after inputting picture L into the second edge detection model, the second edge detection model can determine that the color blocks in picture L include Figure 20 the multiple items identified by the dashed boxes in
[0165] S1811: The edge detection model sends the edge recognition result to the schedule management service.
[0166] As an example, after receiving the edge recognition result, the schedule management service can cache the edge recognition result.
[0167] It is worth mentioning that after performing text recognition and edge recognition processing on picture L through the above two branches respectively, a first recognition result and a second recognition result can be obtained. The first recognition result includes a text recognition result and target indication information, and the second recognition result includes an edge recognition result and a picture category. Since the text recognition result can represent the text content in picture L and the edge recognition result can represent the layout of picture L, subsequent extraction of schedule information based on these two types of data, namely the first recognition result and the second recognition result, can improve the accuracy of information extraction. The specific implementation can be seen in the following steps.
[0168] S1812: When the picture category of picture L is an instant messaging chat screenshot, the schedule management service filters the text recognition result and the edge recognition result according to the first filtering rule.
[0169] Different picture categories correspond to different filtering rules. In implementation, the schedule management service determines the corresponding filtering rule according to the picture category of picture L, and then filters the text recognition result and the edge recognition result of picture L according to the determined filtering rule to filter out interference information irrelevant to the schedule.
[0170] In one example, the instant messaging chat screenshot corresponds to the first filtering rule. The first filtering rule may include filtering out the recognition data corresponding to skewed lines, small characters, and chat timestamps respectively. A skewed line refers to a text line corresponding to a chat element with an inclination angle greater than a preset angle, and the preset angle can be set according to requirements. For example, the preset angle is 10 degrees; small characters refer to a text line corresponding to a chat element with a line height less than a target line height, and the target line height may refer to the average line height of all text lines in picture L; a chat timestamp is a timestamp used to indicate the chat time. Optionally, the first filtering rule may also include filtering out dense text blocks, floating text on the picture, "back" in the upper left corner, "+", etc. in the lower right corner. Exemplarily, see Figure 19 , Figure 19 According to an exemplary embodiment, some chat elements in picture L that need to be filtered are identified, including chat timestamp 1101, small characters 1102, skewed line 1103, and "+" 1004.
[0171] As an example, in the implementation of filtering out the recognition data corresponding to skewed lines, the inclination angle of each text line can be calculated based on the text line coordinates in the text recognition result, so as to determine which text lines are skewed lines, and then delete the recognition data corresponding to the skewed lines. For example, delete the text line coordinates and text line recognition content of the skewed lines.
[0172] As an example, in the implementation of filtering the recognition data corresponding to small characters, the line height of each text line in picture L can be determined according to the text line coordinates in the text recognition result, so as to filter the recognition data corresponding to the text lines with a line height less than the target line height. For example, filter the text line coordinates and the text line recognition content of the text lines with a line height less than the target line height.
[0173] As an example, in the implementation of filtering the recognition data corresponding to the chat timestamp, chat elements with the category of chat timestamp can be determined according to the edge recognition result, and at least one first candidate chat element can be obtained. According to the coordinates of each first candidate chat element in at least one first candidate chat element and the text line coordinates in the text recognition result of picture L, match the text line recognition content corresponding to each first candidate chat element from the text recognition result of picture L. When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element, and the target first candidate chat element is any one of at least one first candidate chat element. When it is determined to filter the recognition data corresponding to the target first candidate chat element, filter the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of picture L.
[0174] As can be seen from the foregoing description, the edge recognition result of the instant messaging chat screenshot includes the category of chat elements. Therefore, according to the category of chat elements in the edge recognition result, chat elements with the category of chat timestamp can be filtered out to obtain at least one first candidate chat element. Since the edge recognition result does not include the text line recognition content, that is, it is impossible to know the text content corresponding to each chat element. In some possible cases, the edge detection model may misjudge the category of a chat element that is not a chat timestamp as a chat timestamp. Therefore, in order to avoid filtering out chat elements that are not chat timestamps as much as possible, after determining at least one first candidate chat element, the text line recognition content corresponding to each first candidate chat element can be matched from the text recognition result of picture L according to the coordinates of each first candidate chat element and the text line coordinates of picture L. For example, for any one of the first candidate chat elements, according to the coordinates of this first candidate chat element and the text line coordinates of picture L, the text line recognition content of at least one text line located in the area corresponding to this first candidate chat element is determined from the text recognition result of picture L, so as to match the text line recognition content corresponding to this first candidate chat element. Then, the schedule management service can determine whether the matched text line recognition content is a timestamp through the target natural language model. For example, the schedule management service can send the matched text line recognition content to the target natural language model through getEntitiy() to request the target natural language model to identify whether the text line recognition content is a timestamp. If it is determined through the target natural language model that the text line recognition content is a timestamp, it can be further determined that this first candidate chat element may be a chat timestamp. Otherwise, if it is determined through the target natural language model that the text line recognition content is not a timestamp, it can be determined that this first candidate chat element is not a chat timestamp.
[0175] Since the position of the chat timestamp in the instant messaging chat screenshot is generally fixed. For example, it is usually centered in the chat window of WeChat. Therefore, in the case where it is determined through the target natural language model that a certain first candidate chat element may be a chat timestamp, the coordinates of this first candidate chat element can be further used to determine whether this first candidate chat element is a chat timestamp, so as to improve the accuracy of chat timestamp recognition.
[0176] As an example of the present application, for any one of the at least one first candidate chat elements, the specific implementation of determining whether to filter out the recognition data corresponding to this first candidate chat element according to the coordinates of this first candidate chat element may include the following two cases:
[0177] The first case: When it is determined according to the coordinates of this first candidate chat element that this first candidate chat element is located in the middle position of Picture L and there is a single text line within the area corresponding to this first candidate chat element, determine to filter out the recognition data corresponding to this first candidate chat element.
[0178] Since in the chat windows of most instant messaging applications, the chat timestamp is generally centered and there is only one line of text recognition content within the chat timestamp, that is, only the timestamp. Therefore, if it is determined that this first candidate chat element is located in the middle position of Picture L and there is only a single text line within the area corresponding to this first candidate chat element, and since it has been determined through the target natural language model that the text line recognition content is the timestamp, it can be determined that this first candidate chat element is the chat timestamp, that is, it is determined that the recognition data corresponding to this first candidate chat element can be filtered out.
[0179] The second case: When it is determined according to the coordinates of this first candidate chat element that this first candidate chat element is located on the right side of Picture L and there is a single text line within the area corresponding to this first candidate chat element, if the text line recognition content corresponding to this first candidate chat element only includes time and does not include date, determine to filter out the recognition data corresponding to this first candidate chat element.
[0180] Since in the chat windows of some instant messaging applications (such as the chat interface forwarded in WeChat), the chat timestamp may also be displayed on the right side, and there is only one line of text recognition content within the chat timestamp. Additionally, the chat timestamp only includes time and does not include date. Therefore, if it is determined that this first candidate chat element is located on the right side of Picture L and there is only a single text line within the area corresponding to this first candidate chat element, it can be judged whether the text line recognition content corresponding to this first candidate chat element includes date. For example, it can be determined whether the text line recognition content includes date through the target natural language model. If it is determined that it does not include date, it can be determined that this first candidate chat element is the chat timestamp, that is, it is determined that the recognition data corresponding to this first candidate chat element can be filtered out. Of course, if it is determined that it includes date, it can be determined not to filter.
[0181] In one example, for the second case, it is also possible not to judge whether the text line recognition content corresponding to this first candidate chat element only includes time and does not include date. As long as it is determined that this first candidate chat element is located on the right side of Picture L and there is a single text line within the area corresponding to this first candidate chat element, the schedule management service can determine to filter out the recognition data corresponding to this first candidate chat element.
[0182] When it is determined through the above process that a certain first candidate chat element is a chat timestamp, the schedule management service deletes the recognition data corresponding to this first candidate chat element from the text recognition result and the edge recognition result of picture L. For example, it deletes the text line coordinates and the text line recognition content corresponding to this first candidate chat element from the text recognition result of picture L, and deletes the coordinates and category corresponding to this first candidate chat element from the edge recognition result of picture L. Of course, if it is determined through the above process that a certain first candidate chat element is not a chat timestamp, the schedule management service does not filter out the recognition data corresponding to this first candidate chat element.
[0183] It is worth mentioning that first determining the chat elements whose category is a chat timestamp according to the edge recognition result, then matching the corresponding text line recognition content from the text recognition result, determining whether it is a chat timestamp through the target natural language model based on the matched text line recognition content, and then determining whether it is a chat timestamp according to the position of the chat element can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0184] It should be noted that the above first filtering rule is only exemplary. When the instant messaging chat screenshots are from different instant messaging applications, the arrangement layouts of their chat elements are usually different, and the chat elements may also be different, so that the interference information in the instant messaging chat screenshots of different instant messaging applications may be different. Exemplarily, it usually includes several possible situations shown in Table 4. Therefore, in another example, the first filtering rule may also include other rules for filtering out information unrelated to the chat content.
[0185] Table 4
[0186]
[0187] In order to effectively filter out the interference information in instant messaging chat screenshots from different instant messaging applications, a first filtering rule can be set according to the union of the possible interference information shown in Table 4, so as to ensure that no matter which instant messaging chat screenshot is processed, the interference information can be effectively removed. Exemplarily, the first filtering rule can also include filtering out the text in the avatar, the user name, the specified identifier, etc. In implementation, after filtering out the recognition data corresponding to the chat timestamp in the picture L from the text recognition result and the edge recognition result of the picture L, the schedule management service can determine the chat elements irrelevant to the chat content from the remaining chat elements in the edge recognition result according to the category of each remaining chat element in the edge recognition result, obtain at least one second candidate chat element, and match the text line recognition content corresponding to each second candidate chat element in the text recognition result of the picture L according to the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition result of the picture L. Delete the matched text line recognition content and the corresponding text line coordinates from the text recognition result of the picture L.
[0188] Exemplarily, in the implementation of filtering out the text in the avatar, the schedule management service can determine the chat element whose category is the avatar according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the text recognition result of the picture L according to the coordinates of the chat element. If there is a matched text line recognition content, delete the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content from the text line recognition result of the picture L, so as to delete the text in the avatar.
[0189] Exemplarily, in the implementation of filtering out the user name, the schedule management service can determine the chat element whose category is the user name according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the text recognition result of the picture L according to the coordinates of the chat element. If there is a matched text line recognition content, delete the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content from the text line recognition result of the picture L, so as to delete the user name.
[0190] Exemplarily, in the implementation of filtering out the specified identifier, the schedule management service can determine the chat element whose category is the specified identifier according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the edge recognition result of the picture L. If the text line recognition content is the specified identifier, for example, it is "+", the recognition data corresponding to the chat element can be deleted from the text line recognition result of the picture L, for example, the text line coordinates and the text line recognition content corresponding to the chat element are deleted. The schedule management service can also delete the recognition data corresponding to the chat element from the edge recognition result, for example, delete the coordinates and the category of the chat element.
[0191] It should be noted that the above description is given by taking the picture L as a screenshot of an instant messaging chat as an example. In another example, if the picture L is not a screenshot of an instant messaging chat, for example, it is a screenshot of an instant messaging notification card or an order screenshot, the schedule management service filters out interference information based on the second filtering rule. In one example, in the implementation of filtering based on the second filtering rule, the schedule management service can match the text line recognition content in each color block from the text recognition result of the picture L according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the picture L. For any one of the color blocks, if it is determined according to the text line recognition content in any one of the color blocks that the content related to the schedule is not included in any one of the color blocks, for example, the time information and location are not included, the recognition data corresponding to any one of the color blocks is filtered out from the edge recognition result and the text recognition result of the picture L. In one example, it can be determined whether the color block includes time information and location through the target natural language model. For example, the text line recognition content in the color block can be sent to the target natural language model one by one to request the target natural language model to determine whether the time information and location are included.
[0192] In the case where the picture L is a screenshot of an instant messaging notification card, it can be seen from Table 4 that the possible interference information includes small characters, skewed lines, and a middle timestamp. Therefore, in one example, before filtering based on the second filtering rule, the recognition data corresponding to the skewed lines, small characters, and the middle timestamp in the text line recognition content of the picture L and the edge recognition result can also be filtered out first, and then filtered out based on the second filtering rule. The filtering of the skewed lines, small characters, and the middle timestamp can refer to the first filtering rule.
[0193] As an example of the present application, before filtering, it can also be queried whether the number of color blocks in the picture L is less than the number threshold. If the number of color blocks in the picture L is less than the number threshold, it means that there are no a large number of color blocks in the picture L. In this case, the electronic device can usually process the picture L. Therefore, filtering processing can be performed according to the second filtering rule. If the number of color blocks in the picture L is greater than or equal to the number threshold, it means that the picture L includes a large number of color blocks. In this case, instead of performing filtering processing, a prompt message can be displayed to guide the user to re-take a screenshot of the picture L, for example, guiding the user to intercept a partial area including the schedule information in the picture L through the electronic device. The number threshold can be set according to requirements. For example, the number threshold can be 10.
[0194] As an example of this application, before filtering, the schedule management service can also determine whether there are dense text blocks based on the text block coordinates in the text recognition result of picture L. If there are dense text blocks, a prompt message can be displayed in the calendar application. The prompt message is used to prompt the user that there are dense text blocks, so that the user can re-crop picture L according to the requirements. If there are no dense text blocks, the schedule management service performs the filtering operation.
[0195] Or in another example, during the filtering process, if the schedule management service determines that there are dense text blocks based on the text block coordinates in the text recognition result of picture L, the text line coordinates and text line recognition content in these text blocks are deleted.
[0196] It is worth mentioning that if no filtering is performed, it is easy to cause the subsequent target natural language model to be unable to accurately extract schedule information. For example, for Figure 17 picture p1 in, the time may be extracted as the chat timestamp "20:37", and the address may be extracted as "No. 5 Ningyuan Road, Danling Street, Haidian District, Beijing", etc. In the embodiments of this application, after determining the first recognition result and the second recognition result, filtering out the interference information in picture L can improve the accuracy of subsequent schedule information extraction.
[0197] As an example rather than a limitation, when picture L is a captured picture, since the captured picture may be skewed and include a background, etc., no filtering may be performed, and the text recognition result is used for text splicing later.
[0198] In another example, when picture L is a captured picture, the text recognition result and edge recognition result of picture L can also be filtered according to the second filtering rule. The embodiments of this application do not limit this.
[0199] S1813: The schedule management service matches the text line recognition content of each remaining chat element from the filtered text recognition result based on the filtered edge recognition result.
[0200] The schedule management service matches the text line recognition content corresponding to each filtered chat element from the filtered text recognition result according to the coordinates of each chat element in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.
[0201] When picture L is not an instant messaging chat screenshot, the schedule management service matches the text line recognition content in each filtered color block from the filtered text recognition result according to the coordinates of each color block in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.
[0202] S1814: The schedule management service determines the scene corresponding to picture L based on the text line recognition content of each remaining chat element and the picture category of picture L.
[0203] The scene corresponding to picture L refers to the scene involved in the text content in picture L.
[0204] In the case where it is determined that picture L is an instant messaging chat screenshot based on the picture category of picture L, if it is further determined to be a chat conversation according to the text line recognition content of each filtered chat element, the scene corresponding to picture L can be determined as the IM chat scene.
[0205] In addition, in the case where the picture category of picture L is other pictures, such as an instant messaging notification card screenshot or an order screenshot, the schedule management service can also combine the text line recognition content of each matched color block to determine the scene corresponding to picture L, such as the service notification card scene or the high-speed rail travel order scene.
[0206] S1815: The schedule management service performs chat conversation splicing on the text line recognition content of each filtered chat element based on the scene corresponding to picture L to obtain the target text information.
[0207] In implementation, the schedule management service performs processing such as line breaks, carriage returns, and splicing on the text line recognition content of each filtered chat element according to the scene corresponding to picture L. Since chat content is usually relatively casual, for example, a complete sentence may be sent in multiple messages. In the case where it is determined that the scene corresponding to picture L is the IM chat scene, the schedule management service can splice these multiple messages into one sentence. Therefore, performing chat conversation splicing based on the scene corresponding to picture L can make the spliced text closer to natural language text. By way of example and not limitation, the target text information after chat conversation splicing is consistent with the content displayed in the smeared interface.
[0208] In one example, the spliced chat conversation content includes the conversation type, title, nickname, conversation content, etc., and the nickname can be customized. For example, the chat conversation content can be spliced in the following format:
[0209] Conversation topic: IM chat
[0210] Title: xx group chat
[0211] Nickname A: xxx
[0212] Nickname B: xxx
[0213] Nickname A: xxxx .......
[0215] It should be noted that S1814 to S1816 are optional operations. In another example, the schedule management service can also splice chat conversations based on the text line recognition content of each filtered chat element according to the picture category of picture L.
[0216] In addition, when picture L is other pictures, such as a screenshot of an instant messaging notification card or an order screenshot, the schedule management service splices the text line recognition content in each filtered color block according to the scenario corresponding to picture L, and its implementation can refer to the splicing of chat conversations.
[0217] S1816: The schedule management service constructs a prompt according to the scenario corresponding to picture L and the target text information.
[0218] Pictures of different picture categories correspond to different prompt construction templates. When constructing a prompt, the schedule management service can use the prompt construction template corresponding to the picture category of picture L to construct the prompt. As an example of this application, the constructed prompt includes scenario description information, and the scenario description information is used to indicate the scenario corresponding to picture L, that is, when constructing the prompt, the schedule management service adds the scenario description information of picture L.
[0219] It is worth noting that if the scenario description information is not added when constructing the prompt, then when the target natural language model performs recognition later, the target natural language model is likely to extract each time and the verbs related before and after that time into a schedule information, so it is easy to extract multiple schedule information, resulting in inaccurate extraction of schedule information. Therefore, in order to enable the target natural language model to accurately recognize the schedule information, the schedule management service adds scenario description information to the constructed prompt, so that the target natural language model can extract an accurate schedule information.
[0220] Exemplarily, taking picture L as an instant messaging chat screenshot as an example, the prompt constructed by the schedule management service can be:
[0221] <|Human|>The following content may be an IM chat conversation. The IM chat conversation contains <title, time, location, participants> fields. Each event's existing fields are output in one line in json format without additional reply.\n\nKaiTan
[0222] HarmonyOS Architecture Evolution and Key Technologies
[0223] HDC Together
[0224] Time 09:00 - 16:30, October 23rd
[0225] Location Microsoft Asia-Pacific R & D Group Building
[0226] No. 5, Danling Street, Haidian District, Beijing
[0227] <|Moss|>
[0228] S1817: The schedule management service sends a prompt to the target natural language model.
[0229] After the schedule management service constructs the prompt, it invokes the target natural language model and sends the constructed prompt to the target natural language model for recognition to extract schedule information. Exemplarily, the schedule management service can request the target natural language model to extract schedule information through extractinformation().
[0230] It should be noted that the embodiments of the present application are described by taking the target natural language model deployed in an electronic device as an example. In another example, the target natural language model can also be deployed in the cloud, and the cloud can provide an interface for the electronic device to call the target natural language model. Thus, when the target natural language model is needed, the schedule management service can call the target natural language model through the provided interface. The embodiments of the present application do not limit this.
[0231] S1818: The target natural language model determines schedule information based on the prompt.
[0232] In one example, the schedule information output by the target natural language model includes: Title: KaiTan, Start Time: October 23, 09:00, End Time: 16:30, Location: Microsoft Asia-Pacific R&D Group Building, No. 5, Danling Street, Haidian District, Beijing, Repeat Cycle: (None).
[0233] S1819: The target natural language model sends the schedule information to the schedule management service.
[0234] Exemplarily, the target natural language model can send the schedule information to the schedule management service in JSON format.
[0235] S1820: The schedule management service decodes the received schedule information.
[0236] The main purpose of this step is to split the time information in the received schedule information. For the specific implementation, reference can be made to the foregoing.
[0237] S1821: The schedule management service performs post-processing on the schedule fields of the decoded schedule information.
[0238] For the post-processing of the schedule fields of the schedule information, reference can be made to the foregoing.
[0239] In one example, before creating the schedule information, the schedule management service may also call the personal behavior feature model to request a query of the user's historical behavior data. For example, the historical behavior data includes historical locations, etc. Accordingly, the personal behavior feature model returns the historical behavior data. In this way, the schedule management service can predict the places the user may go based on the historical behavior data, and then combine the schedule information fed back by the target natural language model to create the final schedule information. For example, the predicted address information is added to the schedule information.
[0240] S1822: The schedule management service displays the processed schedule information in the calendar application.
[0241] Exemplarily, when the picture L is a screenshot of an instant messaging chat, the schedule information displayed on the electronic device is as Figure 3 shown at 13 in figure (d) as shown.
[0242] In one example, before displaying the schedule information, a request confirmation notice may be displayed first. After receiving the confirmation display instruction triggered by the user based on the request confirmation notice, the schedule management service then displays the schedule information in the calendar application.
[0243] As an example of the present application, the electronic device also supports the user to edit the displayed schedule information. Exemplarily, it supports the user to modify the title of the schedule information. For example, refer to the Figure 4 embodiment shown. During this process, when the user triggers the electronic device to display the target interface, text that can be smeared needs to be displayed in the target interface. For this purpose, the schedule management service may also request the target natural language model to perform word combination processing on the spliced text. For example, the two words "I" and "we" are combined into "we" so that the text can be displayed in the target interface according to the combined words, thereby facilitating the user to smear. In practice, the schedule management service can call the target natural language model through getWordSegment() to request the target natural language model to perform word combination processing. In addition, the schedule management service can also call the target natural language model through getWordSegment() to request the target natural language model to determine the theme of the spliced text. For example, the target natural language model is specified to extract the theme entity (such as meeting, dinner). In this way, the schedule management service can display the theme in the target interface.
[0244] In an embodiment of the present application, in response to a schedule extraction operation, target text information is obtained, where the target text information includes the schedule information to be extracted. The target natural language model that has been pre-trained is used to perform schedule extraction on the target text information to obtain first schedule information. The target natural language model can extract the schedule information in the text information and combine and represent the time information in the text information. The time information in the first schedule information is split to obtain second schedule information, where the second schedule information includes the time information obtained after splitting and other schedule element information (such as the subject) in the first schedule information except for the time information. The second schedule information is displayed. In this way, it is not necessary for the user to manually input the schedule information item by item in the calendar application, which improves the efficiency of creating schedule information.
[0245] In addition, since the target natural language model combines and represents the time information, it is not necessary to pay attention to many time-related fields, thereby reducing the computational difficulty of the target natural language model. Since the computational difficulty is reduced, the accuracy of schedule information extraction can be improved. Moreover, the length of the schedule information output by the target natural language model is reduced. Since the target natural language model is iterative reasoning and each output result depends on the results of the previous several outputs, when the length of the schedule information is reduced, the processing duration of the target natural language model will be reduced, thereby reducing the latency of schedule information extraction. In addition, the target natural language model does not need to calculate the result in one step. For example, it only needs to infer the start time point and the duration description information. In this way, it is possible to prevent as much as possible the problem of reasoning errors caused by the weak computational ability of most language models, and improve the accuracy of reasoning. Of course, in a possible implementation, if a model with stronger reasoning ability is used for calculation, the reasoning can also be completed in one step.
[0246] It should be further noted that the embodiment of the present application is described by taking the input of the target natural language model as text information as an example. In another example, the input of the target natural language model can also be a picture, that is, the picture can be directly input into the target natural language model for processing so that the target natural language model extracts the schedule information in the picture. In this case, the training data used in the stage of training the target natural language model can include picture training samples and corresponding schedule information samples.
[0247] The electronic device involved in the embodiment of the present application can be a mobile phone, a sports camera (GoPro), a digital camera, a tablet computer, a desktop type, a laptop, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, etc., and the embodiment of the present application does not make a limitation in this regard.
[0248] Figure 21This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Refer to Figure 21 , the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0249] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0250] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0251] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.
[0252] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0253] It can be understood that the interface connection relationships shown among the various modules in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0254] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, the wireless communication module 160, etc.
[0255] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.
[0256] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0257] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is an integer greater than 1.
[0258] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0259] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0260] The internal memory 121 can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0261] The electronic device 100 can implement audio functions, such as music playback, recording, etc., through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.
[0262] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 together form a touch screen, also known as the "touch display screen". The touch sensor 180K is used to detect touch operations acting thereon or in its vicinity. The touch sensor 180K can transmit the detected touch operations to the application processor to determine the type of touch event. Visual outputs related to the touch operations can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a position different from that of the display screen 194.
[0263] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.
[0264] The above are the optional embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the technical scope disclosed in the present application shall be included within the protection scope of the present application.
Claims
1. A method for displaying schedule information, characterized in that, The method includes: In response to a schedule extraction operation, obtaining target text information, where the target text information includes schedule information to be extracted; Performing schedule extraction on the target text information through a pre-trained target natural language model to obtain first schedule information, where the target natural language model can extract schedule information in the text information and represent the time information in the text information in a combined manner; Splitting the time information in the first schedule information to obtain second schedule information, where the second schedule information includes the split time information and other schedule element information in the first schedule information except for the time information; Displaying the second schedule information.
2. The method according to claim 1, characterized in that, The first schedule information includes field values corresponding to at least one of multiple fields, where the multiple fields include a title field, a start time field, an end time field, a location field, and a repeat period field. The field value corresponding to the title field is a title, the field value corresponding to the start time field includes one or more of a start date, a repeat start date, and a start time point, the field value corresponding to the end time field includes one or more of an end date, a repeat end date, and an end time point, and the field value corresponding to the repeat period field is a schedule repeat period.
3. The method according to claim 2, wherein When the target text information includes the start time point and duration description information but does not include the end time point, the field value corresponding to the end time field further includes the duration description information, and the end time point in the first schedule information is represented by the start time point and the duration description information in the target text information, where the duration description information is information for describing the schedule duration.
4. The method according to claim 2 or 3, characterized in that, The field value corresponding to the end time field of the first schedule information includes duration description information, where the duration description information is information for describing the schedule duration; The splitting the time information in the first schedule information to obtain second schedule information includes: Splitting out the start date from the field value corresponding to the start time field of the first schedule information as the field value corresponding to the first date field; Splitting out the end date from the field value corresponding to the end time field of the first schedule information as the field value corresponding to the second date field, where the first date field and the second date field are both newly added fields and are different; Determining the field value corresponding to the split end time field according to the start time point and the duration description information in the field value corresponding to the end time field of the first schedule information.
5. The method according to claim 4, characterized in that, When the first schedule information does not include periodic schedule information, the first date field is the start date field, and the second date field is the end date field; when the first schedule information includes periodic schedule information, the first date field is the repeat start date field, and the second date field is the repeat end date field.
6. The method according to claim 4 or 5, characterized in that Determining the field value corresponding to the end time field after splitting according to the start time point and the duration description information in the field value corresponding to the end time field of the first schedule information includes: Determining the result of adding the start time point in the field value corresponding to the end time field to the duration described by the duration description information as the field value corresponding to the end time field after splitting.
7. The method according to any one of claims 1-5, characterized in that, The target natural language model is obtained by training an initial natural language understanding model based on a sample training set. The sample training set includes multiple groups of sample training data. Each group of sample training data includes a text training sample and a schedule information sample corresponding to the text training sample, and the time information in the schedule information sample is processed by merging.
8. The method according to any one of claims 1-7, characterized in that, The obtaining of the target text information in response to the schedule extraction operation includes: In response to the schedule extraction operation on the first picture, through the target recognition model, determining a first recognition result and a second recognition result. The first recognition result includes the text recognition result of the first picture, and the second recognition result includes the edge recognition result and the picture category of the first picture. The edge recognition result includes the graphic and text element attribute information or the color block attribute information of the first picture. The target recognition model can determine the text recognition result, the edge recognition result and the picture category of any picture; Filtering out the interference information irrelevant to the schedule in the text recognition result and the edge recognition result of the first picture according to the picture category of the first picture; Based on the filtered first recognition result and the filtered second recognition result, obtaining the target text information.
9. The method according to claim 8, characterized in that, The picture category includes instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots and other category pictures. The instant messaging chat screenshot refers to a picture obtained by taking a screenshot of the chat interface in the instant messaging application. The instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of the service notification card in the instant messaging application. The order screenshot refers to a picture obtained by taking a screenshot of the order interface in the application program. The other category pictures include other pictures except the instant messaging chat screenshots, the instant messaging notification card screenshots and the order screenshots; Wherein, when the first picture is the instant messaging chat screenshot, the edge recognition result includes the graphic and text element attribute information. When the first picture is one of the instant messaging notification card screenshot, the order screenshot and the other category pictures, the edge recognition result includes the color block attribute information.
10. The method according to claim 9, characterized in that, The text recognition result includes text line coordinates; Filtering out the interference information irrelevant to the schedule in the text recognition result and the edge recognition result of the first picture according to the picture category of the first picture includes: When it is determined that the first picture is the instant messaging chat screenshot according to the picture category of the first picture, filtering out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture according to the text line coordinates of each text line in the text recognition result of the first picture; Filter out the recognition data corresponding to the text lines with line heights less than the target line height in the text recognition result of the first picture and the edge recognition result according to the text line coordinates of each text line in the text recognition result of the first picture, where the target line height is the average line height of all text lines in the first picture; Filter out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result of the first picture and the edge recognition result.
11. The method according to claim 10, wherein The graphic element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result also includes text line recognition content; The filtering out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result of the first picture and the edge recognition result includes: Determine the chat elements with the category of chat timestamp from the edge recognition result to obtain at least one first candidate chat element; According to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each first candidate chat element from the text recognition result of the first picture; When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element, where the target first candidate chat element is any one of the at least one first candidate chat elements; When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filter out the recognition data corresponding to the target first candidate chat element from the text recognition result of the first picture and the edge recognition result.
12. The method according to claim 11, wherein The determining whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element includes: When it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located in the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, determine to filter out the recognition data corresponding to the target first candidate chat element; or, When it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located in the right position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes a time point and does not include a date, determine to filter out the recognition data corresponding to the target first candidate chat element.
13. The method according to claim 9, wherein The text recognition result includes text line coordinates and text line recognition content; The filtering out the interference information irrelevant to the schedule from the text recognition result of the first picture and the edge recognition result according to the picture category of the first picture includes: In the case where it is determined that the first picture is a screenshot of the instant messaging notification card or an order screenshot according to the category of the first picture, match the text line recognition content in each color block from the text recognition result of the first picture according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture; For any one of the color blocks, if it is determined that the content related to the schedule is not included in the color block according to the text line recognition content in the color block, filter out the recognition data corresponding to the color block from the edge recognition result and the text recognition result of the first picture.
14. The method according to any one of claims 8 - 13, characterized in that, The text recognition result includes text line coordinates and text line recognition content; The obtaining of the target text information based on the filtered first recognition result and the filtered second recognition result includes: Based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each object from the text recognition result of the filtered first picture, where the object is a chat element or a color block; Determine the scene corresponding to the first picture according to the text line recognition content corresponding to each object and the picture category of the first picture; Based on the scene corresponding to the first picture, perform text splicing on the text line recognition content corresponding to each object to obtain the target text information.
15. The method according to claim 14, wherein After displaying the second schedule information, it further includes: In response to an editing operation on the title of the second schedule information, display a target interface, and the target interface includes the target text information; In response to a selection operation on the text line content displayed in the target interface, input the text selected by the selection operation into the title input box; In response to an end-of-editing operation, modify the title to the content input in the title input box.
16. The method according to any one of claims 8-15, characterized in that, The obtaining of the first schedule information by performing schedule extraction on the target text information through a pre-trained target natural language model includes: According to the target text information and the scene corresponding to the first picture, construct a prompt, where the prompt includes scene description information for describing the scene corresponding to the first picture, and the scene corresponding to the first picture is determined based on the first recognition result and the second recognition result; Input the prompt into the target natural language model for processing to obtain the first schedule information.
17. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of claims 1-16 is implemented.
18. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when running on a computer, cause the computer to execute the method described in any one of claims 1-16.
Citation Information
Patent Citations
Method and device for generating calendar reminding
CN104463552A
Schedule processing method, device and system
CN111182154A
Schedule information synchronization method and device and electronic equipment
CN115018442A
Calendar view display method, electronic equipment and readable storage medium
CN115079915A
Target information processing method and apparatus, and electronic device
WO2023221895A1
Cited By
Automatic schedule acquisition method for mobile operating system
CN121999475A