Method for acquiring schedule information, electronic equipment and readable storage medium
Through the system prompt KV value and image recognition technology of cached natural language model, the problem of low efficiency in creating agenda information in electronic devices is solved, and fast and accurate extraction of agenda information is achieved.
Patent Information
- Application Number
- CN202410028265.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-01-05
AI Technical Summary
In the prior art, electronic devices are inefficient in creating schedule information in third-party social applications, and require cumbersome manual operations from users.
The KV value of the system prompt is cached by the pre-trained natural language model, reducing the amount of real-time calculation, and quickly obtaining the agenda information of the target text information. Combining picture recognition and edge detection, interfering information is filtered, and the speed and accuracy of agenda information extraction are improved.
It realizes the rapid and accurate creation of schedule information in electronic devices, reduces user manual operations, and improves the efficiency and accuracy of schedule information creation.
Smart Images

Figure CN120338735A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and particularly to a method for obtaining schedule information, an electronic device, and a readable storage medium. Background Art
[0002] With the rapid development of terminal technology, electronic devices can usually install various types of third-party social applications, such as instant messaging applications, ticket purchasing applications, etc. Some messages or service notifications in third-party social applications often involve schedule information. For example, the conversation content in the chat window of an instant messaging application involves the schedule information of a certain meeting. In some scenarios, users usually need to create the schedule information involved in third-party social applications through an electronic device. Therefore, how to enable an electronic device to quickly create schedule information has become a research hotspot in this field. Summary of the Invention
[0003] This application provides a method for obtaining schedule information, an electronic device, and a readable storage medium, which can solve the problem of how to quickly create schedule information in related technologies. The technical solutions are as follows:
[0004] In a first aspect, a method for obtaining schedule information is provided, which is applied to an electronic device. The method includes:
[0005] In response to a schedule extraction operation, obtain target text information, where the target text information includes the schedule information to be extracted. Determine the system prompt corresponding to the target text information from multiple system prompts to obtain a target system prompt. The system prompt is used to describe the task of the event to be inferred, and each system prompt in the multiple system prompts corresponds to a business scenario. Obtain the key-value (KV) value of the target system prompt from the cached KV values of the multiple system prompts. Use the target text information and the KV value of the target system prompt as the input of a pre-trained target natural language model, and perform inference through the target natural language model to output the schedule information corresponding to the target text information.
[0006] In this way, by pre-computing and caching the KV values corresponding to different system prompts, during the application process, the target system prompt to be used can be determined according to the business scenario, and then the KV value corresponding to the target system prompt is obtained from the cached KV values as a part of the model input, so that the target natural language model does not need to calculate the target system prompt again, which can reduce the calculation amount of the target natural language model, thereby accelerating the extraction speed of schedule information and improving the creation speed of schedule information.
[0007] As an example of the present application, to determine the system prompt corresponding to the target text information from multiple system prompts, it includes: determining the business scenario corresponding to the target text information according to one or more of the keyword, business entry information, application information, and intent description information related to the target text information, where different keywords are used to indicate different business scenarios, the business entry information is used to indicate the entry used to request text information reasoning, the application information is used to indicate the category of the application from which the target text information comes, and the intent description information is used to indicate the intent of the target text information. Obtain the system prompt corresponding to the determined business scenario from multiple system prompts. In this way, by determining the business scenario of the target text information and obtaining the corresponding target system prompt, the target natural language model can accurately reason according to the target system prompt, improving the accuracy of schedule information extraction.
[0008] As an example of the present application, the target text information is from a picture. Correspondingly, to determine the business scenario corresponding to the target text information according to one or more of the keyword, business entry information, application information, and intent description information related to the target text information, it includes: in the case that the first picture corresponding to the target text information is for schedule extraction requested through a calendar application, if the first picture is obtained by taking a screenshot of the foreground application on an electronic device, determine the business scenario corresponding to the target text information according to the application information of the foreground application. If the first picture is not obtained by taking a screenshot on the electronic device, determine the intent description information corresponding to the first picture through a pre-trained intent classification model, and determine the business scenario corresponding to the target text information according to the intent description information corresponding to the first picture. In this way, for the first picture, the corresponding business scenario can be determined according to its source or its intent, so as to determine the business scenario corresponding to the target text information.
[0009] As an example of the present application, before obtaining the KV value of the target system prompt from the KV values of multiple cached system prompts, it further includes: for each system prompt among the multiple system prompts, input each system prompt into the target natural language model for processing to obtain the KV value of each system prompt. Cache the KV values of each system prompt among the multiple system prompts. In this way, the required KV value can be directly obtained from the cached KV values during the application process. In addition, since the model for calculating the KV value corresponding to each system prompt is the target natural language used during the application process, that is, there is no error in the KV value, it can enable the target natural language model to accurately reason during the application process.
[0010] As an example of the present application, in response to a schedule extraction operation on a first picture, a first recognition result and a second recognition result are determined through a target recognition model. The first recognition result includes the text recognition result of the first picture, and the second recognition result includes the edge recognition result and the picture category of the first picture. The edge recognition result includes the graphic element attribute information or the color block attribute information of the first picture. According to the picture category of the first picture, the interference information irrelevant to the schedule in the text recognition result of the first picture and the edge recognition result is filtered out. Based on the filtered first recognition result and the filtered second recognition result, the target text information is obtained. In this way, it is not necessary for the user to manually create schedule information in the calendar application, improving the efficiency of schedule information creation.
[0011] As an example of the present application, the picture category includes instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. An instant messaging chat screenshot refers to a picture obtained by taking a screenshot of the chat interface in an instant messaging application. An instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of the service notification card in an instant messaging application. An order screenshot refers to a picture obtained by taking a screenshot of the order interface in an application. Other category pictures include other pictures except instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots. Among them, when the first picture is an instant messaging chat screenshot, the edge recognition result includes graphic element attribute information. When the first picture is one of an instant messaging notification card screenshot, an order screenshot, and other category pictures, the edge recognition result includes color block attribute information.
[0012] In this way, by classifying the picture category into the above-mentioned multiple types, it is possible to adopt a targeted method for picture processing according to the picture category in the subsequent schedule information extraction process, thereby improving the effectiveness and accuracy of schedule information extraction.
[0013] As an example of the present application, the text recognition result includes text line coordinates. Correspondingly, according to the picture category of the first picture, the specific implementation of filtering out the interference information unrelated to the schedule in the text recognition result and the edge recognition result of the first picture may include: when the first picture is an instant messaging chat screenshot, according to the text line coordinates of each text line in the text recognition result of the first picture, filtering out the recognition data corresponding to the skewed text lines in the text recognition result of the first picture. According to the text line coordinates of each text line in the text recognition result of the first picture, filtering out the recognition data corresponding to the text lines with a line height less than the target line height in the text recognition result and the edge recognition result of the first picture, and the target line height is the average line height of all text lines in the first picture. Filtering out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture. In this way, by filtering out the recognition data corresponding to the skewed lines, small characters, and chat timestamps in the first picture, the interference information unrelated to the schedule can be effectively removed, that is, the information that may interfere with the extraction of schedule information can be removed, thereby improving the accuracy of subsequent schedule extraction.
[0014] As an example of the present application, the graphic and text element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result also includes text line recognition content. Correspondingly, the specific implementation of filtering out the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture may include: determining the chat elements with the category of chat timestamp from the edge recognition result to obtain at least one first candidate chat element. According to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture, matching the text line recognition content corresponding to each first candidate chat element in the text recognition result of the first picture. When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determining whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element, and the target first candidate chat element is any one of the at least one first candidate chat elements. When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filtering out the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of the first picture.
[0015] In this way, first determine the chat elements with the category of chat timestamp according to the edge recognition result, then match the corresponding text line recognition content from the text recognition result, determine whether it is a chat timestamp through the target natural language model according to the matched text line recognition content, and then determine whether it is a chat timestamp according to the position of the chat element, which can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0016] As an example of the present application, the specific implementation of determining whether to filter the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element may include: when it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located in the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, it is determined to filter the recognition data corresponding to the target first candidate chat element. Alternatively, when it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located on the right side of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes a time point and does not include a date, it is determined to filter the recognition data corresponding to the target first candidate chat element. In this way, by judging whether the first candidate chat element is a chat timestamp according to the position feature and text line feature of the chat timestamp, the accuracy of the judgment can be improved.
[0017] As an example of the present application, the text recognition result includes text line coordinates and text line recognition content. The specific implementation of filtering out interference information unrelated to the schedule in the text recognition result and edge recognition result of the first picture according to the picture category of the first picture may include: when the first picture is a screenshot of an instant messaging notification card or an order screenshot, according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content in each color block from the text recognition result of the first picture. For any one of the color blocks, if it is determined according to the text line recognition content in any one of the color blocks that the color block does not include content related to the schedule, filter out the recognition data corresponding to the color block from the edge recognition result and the text recognition result of the first picture. In this way, by filtering out the color blocks that do not include the schedule, it is convenient to only process the color blocks that include content related to the schedule subsequently, which can improve the data processing efficiency and the accuracy of schedule information extraction.
[0018] As an example of the present application, the text recognition result includes the text line coordinates and the recognized content of the text line. Based on the filtered first recognition result and the filtered second recognition result, the specific implementation of obtaining the target text information may include: Based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the recognized content of the text line corresponding to each object from the text recognition result of the filtered first picture, where the object is a chat element or a color block. Determine the scene corresponding to the first picture based on the recognized content of the text line corresponding to each object and the picture category of the first picture. Based on the scene corresponding to the first picture, perform text splicing on the recognized content of the text line corresponding to each object to obtain the target text information. In this way, by determining the scene corresponding to the first picture and performing text splicing according to the scene corresponding to the first picture, the finally obtained target text information can be made to conform to the natural language rules as much as possible.
[0019] In a second aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for displaying schedule information as described in the first aspect above is implemented.
[0020] In a third aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is made to execute the method for displaying schedule information as described in the first aspect above.
[0021] In a fourth aspect, a computer program product containing instructions is provided. When it runs on a computer, the computer is made to execute the method for displaying schedule information as described in the first aspect above.
[0022] The technical effects obtained in the second aspect, the third aspect, and the fourth aspect above are similar to the technical effects obtained by the corresponding technical means in the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a schematic diagram of a model inference process shown according to an exemplary embodiment;
[0024] Figure 2 is a schematic diagram of an application scenario shown according to an exemplary embodiment;
[0025] Figure 3 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0026] Figure 4 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0027] Figure 5 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0028] Figure 6 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0029] Figure 7 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0030] Figure 8 is a schematic diagram of an application scenario shown according to another exemplary embodiment;
[0031] Figure 9 is a schematic diagram of a framework for displaying schedule information shown according to another exemplary embodiment;
[0032] Figure 10 is a schematic diagram of a framework for displaying schedule information shown according to another exemplary embodiment;
[0033] Figure 11 is a schematic diagram of a software system of an electronic device shown according to an exemplary embodiment;
[0034] Figure 12 is a schematic flowchart of a method for displaying schedule information shown according to an exemplary embodiment;
[0035] Figure 13 is a schematic diagram of the processing of an instant messaging chat screenshot shown according to an exemplary embodiment;
[0036] Figure 14 is a schematic diagram of the processing of an instant messaging notification card screenshot shown according to an exemplary embodiment;
[0037] Figure 15 is a schematic diagram of the architecture of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0038] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the implementation manners of the present application in detail with reference to the accompanying drawings.
[0039] It should be understood that the "multiple" mentioned in this application refers to two or more. In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B. The "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solution of this application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit being different.
[0040] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways.
[0041] Before introducing the method for obtaining schedule information provided by the embodiments of this application, the terms or nouns involved in the embodiments of this application will be briefly described first.
[0042] One-stop office software: It can provide a variety of office functions, such as document processing, spreadsheets, presentation production, project management, calendars, and emails. Users can complete multiple office tasks in the same software, improving the work efficiency of users. One-stop office software usually adopts a unified user interface design, making the switching and use between various functional modules more convenient and consistent.
[0043] Global collection: It means that users can slide their fingers up and down on the screen with three fingers to trigger the electronic device to collect information.
[0044] Magic text: It is a function for quickly extracting text from pictures. Usually, users can turn on or off the switch of Magic text through the path "Settings > Smart Assistant > Magic text", thereby turning on or off this function.
[0045] Large Language Model (LLM): It is a model based on machine learning and natural language processing technologies. By training on a large amount of sample training data, it can learn the ability to serve human language understanding and generation. For example, a trained LLM can be used to reason about text information to extract schedule information from it.
[0046] Prompt: It includes user prompt and system prompt. The user prompt is the input text information, that is, the actual question input by the user. The system prompt is a piece of text or instruction added to the user prompt. The system prompt is a task prompt that describes the event to be inferred in the user prompt. Usually, the two parts of the user prompt and the system prompt jointly constitute the input of the LLM, so as to guide the LLM to generate more accurate and targeted outputs. The system prompt can be a question, a description, a formatted input, or even some keywords. By reasonably designing the prompt, the LLM can be guided to understand the intention of the question and thus generate relatively accurate answers. Exemplarily, the system prompt includes: "The following content may contain several events, and each event contains <'title','start time', 'end time','repeat cycle', 'location', 'url link', 'participants'> fields. The fields existing in each event are output in one line in json format without additional replies.\n\n ;", and the user prompt includes: "At 3 pm on Wednesday, I need to go to Xinhua Bookstore to apply for a card and buy some books, and I need to select a short period of time. On Sunday, I need to go to Shanghai General Motors to repair the host and transmission and be responsible for car washing and interior cleaning and maintenance. Well, this matter is very important. On Saturday, I need to do some cleaning, clean the house, and clean the refrigerator, TV, computer, etc. I may be busy at that time." Based on this, after splicing the system prompt and the user prompt and inputting them into the pre-trained LLM for inference, the output result example includes: "[{'title': 'Apply for a card and buy some books','start time': 'Wednesday 03:00 pm', 'location': 'Xinhua Bookstore'}, {'title': 'Clean the house','start time': 'Saturday', 'location': 'home'}, {'title': 'Car washing and interior cleaning and maintenance','start time': 'Sunday', 'location': 'Shanghai General Motors to repair the host and transmission'}]". It can be seen that the system prompt + user prompt is the real input of the LLM, where the system prompt indicates what the task of extracting events is, such as extracting the content of 7 fields including title, start time, etc. and returning them in json format; the user prompt is the string text to be extracted, that is, the input text information; the above result example is the real extracted result, where '[' is the first character generated.
[0047] Vertical domain: It refers to providing specific services for a limited group, including industries such as entertainment, medical care, environmental protection, education, and sports.
[0048] Currently, various third-party social applications can be installed in an electronic device. Exemplarily, they include, but are not limited to, instant messaging (IM) applications, ticket booking applications, etc. For example, the instant messaging applications can be WeChat, QQ, DingTalk, Feishu, etc., and the ticket booking applications can be Ctrip, Qunar, Tongcheng, Fliggy, etc. Users can chat and book tickets through the third-party social applications, and can also follow the official accounts (or notification accounts) of various industries through the third-party social applications to understand the fields and events to be concerned through the service notification cards issued by the official accounts (or notification accounts). In some scenarios, some messages in the third-party social applications may involve schedule information. For example, one or more messages in the chat window of the instant messaging application involve the relevant schedule of an event. Users usually have the need to record schedule information. In this case, generally, users need to manually create schedule information in the calendar application. However, manually creating schedule information is cumbersome, resulting in low efficiency in creating schedule information.
[0049] In some embodiments, to improve the efficiency of creating schedule information, the electronic device automatically extracts schedule information through a pre-trained natural language model (such as LLM). For example, see Figure 1 , in implementation, the electronic device obtains text information including schedule information (i.e., user prompt), and obtains the system prompt, splices the user prompt and the system prompt to obtain an Instruction, and inputs the Instruction into the pre-trained LLM for inference, and the LLM extracts the schedule information. Among them, the inference process of the LLM is iteratively implemented. Exemplarily, such as Figure 1As shown, after the Instruction is input into the LLM, the LLM calculates the attention based on the Instruction to determine the KV values corresponding to the Instruction. As an example, the LLM can cache the KV values corresponding to the Instruction. Then, the LLM determines the hidden layer output vector based on the KV values corresponding to the Instruction. The hidden layer output vector can be used to predict the next character. The LLM calculates the character probability value according to the hidden layer output vector and outputs a token according to the maximum character probability value. A token is the basic unit output or processed by the LLM. A token can be a word or a character, and is usually also referred to as a character. After that, the electronic device concatenates the first output token to the end of the Instruction to reconstruct the model input data, and inputs the reconstructed input data into the LLM to continue the second inference. Since the LLM cached the KV values corresponding to the Instruction during the first inference, during the second inference, the LLM can only calculate the KV values corresponding to the newly added token in the input data, and based on the cached KV values and the KV values corresponding to the newly added token, determine the hidden layer output vector again, and output the second token based on the hidden layer output vector determined again. And so on, the electronic device concatenates the token output in the nth time to the end of the input data input into the LLM in the nth time to reconstruct the input data used for the (n + 1)th inference. For example, the input data used for the (n + 1)th inference is Instruction, N1, N2, N3...N i ...N n , where N i is the result output in the ith inference. Then, the LLM performs the (n + 1)th inference based on the reconstructed input data until the LLM outputs the termination symbol to end the inference. At this time, all the output results (i.e., all the tokens inferred) are determined as the schedule information to be extracted. Among them, the termination symbol is a symbol predefined during the training process. When this symbol is inferred, it represents the end of the inference. As an example but not a limitation, the termination symbol can be the "EOS" symbol.
[0050] In the above embodiments, the inference latency of the LLM includes the inference duration of the first token and the sum of the inference durations of all subsequent tokens (non-first tokens). The inference duration of the first token includes the inference duration of the user prompt part and the inference duration of the system prompt part. Since the length of the system prompt is usually long, with an average length generally reaching 100 tokens, this results in a fixed latency overhead for extracting schedule information through the LLM, while the computing power and power consumption of the end device are limited. Therefore, the embodiments of the present application provide a method for obtaining schedule information, which can enable the electronic device to automatically create schedule information, improve the creation efficiency of schedule information, and moreover, this method pre-computes the KV values of the system prompt to reduce the amount of data that the LLM needs to compute in real time, so as to reduce the inference duration of the LLM, thereby reducing the latency for obtaining schedule information and improving the speed of obtaining schedule information. Its specific implementation can be seen below.
[0051] As an example of the present application, referring to Table 1, the electronic device can create the schedules involved in the following scenarios into schedule information:
[0052] Table 1
[0053]
[0054] The above graphics and texts include pictures and text information, that is, the method provided by the embodiments of the present application can not only automatically extract the schedule information in the text information, but also automatically extract the schedule information in the pictures. Specifically, the scope supported by the method provided by the embodiments of the present application is shown in Table 2:
[0055] Table 2
[0056]
[0057] According to Table 2, the electronic device can extract the schedule information in the pictures. The pictures can be screenshot pictures or pictures taken by a camera. The form of the screenshot pictures can include but is not limited to full-screen screenshots or area screenshots. The windows involved in the screenshot pictures can be full-screen, split-screen or floating windows. The sources of the screenshot pictures usually come from mobile phones, tablets, PCs, etc.; the form of the pictures taken by the camera can be but is not limited to printed text, handwritten text and artistic words. In addition, the electronic device can also extract the schedule information in the text. The form of the text includes but is not limited to Text and Webview. The format of the text can be plain text, formatted text or text mixed with pictures. The length of the text can be single-paragraph, multi-paragraph, short text or long text, etc.
[0058] In addition, the electronic device can also extract schedule information from the voice. For example, it can receive the voice through the voice assistant and then convert it into text information for processing. In the embodiments of the present application, the input is mainly taken as an example of a picture for illustration.
[0059] For the convenience of understanding, the application scenarios provided by the embodiments of the present application will be introduced next.
[0060] In one example, the multi-round conversations in the chat window of WeChat have a schedule intention. When the user wants to create relevant schedule information, the user can trigger the mobile phone to take a screenshot of the application interface where the chat window is located. For example, the user can double-click on the screenshot trigger area of the mobile phone screen to trigger the screenshot operation, so that the mobile phone takes a screenshot of the interface where the WeChat chat window is located. After that, referring to Figure 2 Figure (a) in, the mobile phone displays the screenshot editing interface U1 and displays the instant messaging chat screenshot obtained after the screenshot in the screenshot editing interface U1, that is, displays the picture p1. The screenshot editing interface U1 includes a "Share" control. When the user wants to create the schedule information in the picture p1, the user can click the "Share" control. Referring to Figure 2 Figure (b) in, in response to the user's trigger operation on the "Share" control, the mobile phone displays the sharing floating window 10. The sharing floating window 10 includes a calendar icon 11, and the user can click the calendar icon 11. In response to the user's trigger operation on the calendar icon 11, the mobile phone starts to process the picture p1 to extract and create relevant schedule information. Exemplarily, referring to Figure 2 Figure (c) in, during this process, the mobile phone can display a prompt message of "Parsing schedule information offline" so that the user can know that the picture p1 is currently being processed for schedule extraction. Referring to Figure 2 Figure (d) in, after the schedule information is successfully created, the mobile phone displays the schedule display interface U2. The schedule display interface U2 includes a schedule display window 12, and the created schedule information 13 is displayed in the schedule display window 12. The schedule information 13 includes content such as a title, a start time point, an end time point, a start date, a location, etc. In this way, the mobile phone realizes the purpose of automatically creating and displaying schedule information, which can avoid the need for the user to manually record and improve the efficiency of creating schedule information.
[0061] As an example of the present application, the schedule display window 12 also includes a plurality of editing controls, and the user can also edit the schedule information 13 created by the mobile phone based on the plurality of editing controls according to the needs. For example, referring to Figure 3 Figure (a) in, when the user wants to modify the title of the schedule, the user can click the title editing control 14 in the schedule display window 12. In response to the user's click operation on the title editing control 14, the mobile phone displays as Figure 3The target interface U3 (which can be called the smearing interface) shown in Figure (b) therein. Information related to the schedule in picture p1 is displayed in the target interface U3. In this way, the user can modify or fill in the title of the schedule by smearing on the information displayed in the target interface U3. Correspondingly, the mobile phone inputs the content smeared by the user in the title input box 15 of the target interface U3. For example, refer to Figure 3 Figure (b) therein. When the user wants to modify the title of the schedule to "HarmonyOS Talk", the user can smear the content of "HarmonyOS", "Talk" in sequence in the target interface U3. Correspondingly, the mobile phone inputs "HarmonyOS", "Talk" in sequence in the title input box 15. Refer to Figure 3 Figure (c) therein. After the user finishes smearing, the user can trigger the "Input" control (or "√" control) of the target interface U3. In response to the triggering operation of the user on the "Input" control (or "√" control), the mobile phone resumes displaying the schedule display window 12. At this time, the user can see that the title of the schedule in the schedule display window 12 is modified to the content modified by the user through the smearing operation, that is, the title of the schedule is changed from "Talk KaiTan" to "HarmonyOS Talk". In this way, by displaying information related to the schedule and smearable in the target interface U3, the user can quickly modify the title of the schedule by smearing, improving the user experience.
[0062] In addition, refer to Figure 3 Figure (a) therein. The schedule display window 12 also includes a time editing control. When the user wants to edit the time information in the schedule information 13, the user can also modify it based on the time editing control. For example, the user can click on the displayed time information to edit it. In addition, after the user swipes down the schedule display window 12, the schedule display window 12 can also provide other editing controls, such as editing controls for the number of repetitions, reminder time, important reminder, etc. In this way, the user can edit the schedule information based on other editing controls, which is not limited in this embodiment of the present application.
[0063] As an example of the present application, after the user clicks the "√" control in the schedule display window 12, in response to this triggering operation, the mobile phone displays the schedule information 13 in the schedule details area of the calendar application, so that the user can view the schedule information 13 from the schedule details area of the calendar application. As an optional example, after the user clicks the "√" control in the schedule display window 12, the mobile phone can also display the schedule information 13 in the form of a card at positions such as the desktop, the negative first screen, or the notification center, etc., for the user to quickly view later, which is not limited in this embodiment of the present application.
[0064] It should be noted that the above is only an example where the user triggers the mobile phone to create schedule information through the sharing entry (i.e., the sharing control). In another example, see Figure 3 In figure (a) of Figure 3 , the mobile phone provides a Magic text control 00 in the screenshot editing interface U1. When the user needs the mobile phone to create schedule information based on the picture p1, the user can also click the Magic text control 00, thereby triggering the mobile phone to create and display schedule information with one click.
[0065] In another example, the user can also trigger the mobile phone to create and display schedule information through the Any Door entry. Exemplarily, see Figure 4 In figure (a) of Figure 4 , after the mobile phone takes a screenshot of the interface where the WeChat chat window is located, the obtained picture p1 is automatically saved to the gallery. Thus, when the user wants the mobile phone to automatically create the schedule information in the picture p1, the user can open the screenshot picture interface U4 in the gallery, and the picture p1 is displayed in the screenshot picture interface U4. See Figure 4 In figure (b) of Figure 4 , the user can trigger the mobile phone to select the picture p1, and then, the user can drag the picture p1 to the right side of the mobile phone screen. When the user drags to a certain position, in response to the user's drag operation, the mobile phone displays application programs that can receive and process the picture p1, such as Figure 4 as shown in figure (b) of Figure 4 , the calendar, WeChat, and QQ are displayed. Thus, the user can continue to drag the picture p1 to the calendar application and then release it. In response to the user's release operation, the mobile phone starts to process the picture p1 to extract and create relevant schedule information. See Figure 4 In figure (c) of Figure 4 , after the mobile phone creates the schedule information, the schedule information is displayed in the schedule display window 12.
[0066] In another example, the user can only select the content related to the schedule in the picture p1, and then trigger the mobile phone to extract the schedule information for the selected part. For example, see Figure 5 In figure (a) of Figure 5 , the user can select a part of the content in the picture p1. For example, a control for triggering the selection operation can be provided in the screenshot editing interface, and after triggering this control, the user can select on the picture p1. In response to the user's selection operation, the mobile phone selects the part selected by the user. As an example, the mobile phone can take a screenshot of the area selected by the user, as Figure 5 shown in figure (b) of Figure 5 . Then, the user can trigger the mobile phone to extract the schedule information for the selected area through an interaction entry such as the "Share" control, and create and display the schedule information related to the content in this area. Thus, by supporting the user to select the picture p1, the data processing amount of the mobile phone can be reduced, and thus the creation efficiency of the schedule information can be improved.
[0067] It should be noted that the interaction entrances used by the user to trigger the mobile phone to extract schedule information from the picture p1 in the above application scenarios are only exemplary. In some embodiments, the mobile phone can also be triggered to extract schedule information from the picture p1 through other interaction entrances. For example, it can also be triggered through interaction entrances such as global collection. The embodiments of the present application do not limit this.
[0068] The above application scenarios are only exemplary. In addition, the mobile phone can also extract schedule information from some types of pictures in other scenarios. Exemplarily, see Figure 6 Figure (a) in. There is a forwarded picture p2 (which can be obtained by screenshot or shooting) in a certain chat window, and there is schedule information in the picture p2. When the user wants to extract the schedule information in the picture p2, the picture p2 can be dragged to the right side of the mobile phone screen. See Figure 6 Figure (b) in. When the user drags to a certain position, in response to the user's drag operation, the mobile phone displays application programs that can receive and process the picture p2. For example, as shown in Figure 6 Figure (b) in, the calendar, WeChat, and QQ are displayed. After the user drags the picture p2 to the calendar application and releases it, in response to the user's release operation, the mobile phone starts to automatically process the picture p2 to extract and create relevant schedule information. See Figure 6 Figure (c) in. After the mobile phone creates the schedule information, the schedule information is displayed in the schedule display window 12.
[0069] See Figure 6 Figure (c) in. In the case where the mobile phone creates schedule information multiple times, multiple schedule information can be displayed in the schedule display window 12. When not all schedule information can be fully displayed in the schedule display window 12, the user can trigger the mobile phone to display the hidden schedule information by swiping the schedule display window 12 left and right.
[0070] In addition, the mobile phone not only supports the user to drag the picture in the chat window to the Any Door entrance, but also supports the user to drag the text to the Any Door entrance. For example, a certain chat window includes chat text, and the chat text includes schedule information. When the user needs the mobile phone to create the schedule information in the chat text, the chat text can be selected, and then the chat text can be dragged to the calendar application in the manner shown in Figure 6 . Accordingly, the mobile phone extracts the schedule information in the chat text, creates and displays the relevant schedule information.
[0071] In another example, the mobile phone can also extract schedule information from the screenshot of the instant messaging notification card. For example, see Figure 7In FIG. (a) below, a schematic diagram of a screenshot of an instant messaging notification card (i.e., picture p3) shown according to an exemplary embodiment is presented. Picture p3 is a screenshot of a notification card published in the service notifications of WeChat and includes schedule information. When a user wants to create and display the schedule information in picture p3 on the mobile phone, they can trigger the mobile phone to extract the schedule information according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information from picture p3 through the sharing entry. Correspondingly, the mobile phone extracts the schedule information based on picture p3 and then creates or displays the relevant schedule information. For example, refer to Figure 7 In FIG. (b) below, the mobile phone creates and displays schedule information 60, which includes a subject, a ticket number, a departure time point, a departure date, a location, etc.
[0072] In another example, the mobile phone can also extract schedule information from a screenshot of an order, where the screenshot of the order can be a screenshot image of a hotel order, a train ticket order, an airplane ticket order, etc. For example, refer to Figure 8 In FIG. (a) below, a schematic diagram of a screenshot of an order (i.e., picture p4) shown according to an exemplary embodiment is presented. Picture p4 is a screenshot of the interface where the train ticket order is located and includes schedule information. When a user wants to create and display the schedule information in picture p4 on the mobile phone, they can trigger the mobile phone according to the operation process described above. For example, the user can trigger the mobile phone to extract the schedule information from picture p4 through the Magic text entry. Correspondingly, the mobile phone extracts the schedule information based on picture p4 and then creates or displays the relevant schedule information. For example, the displayed schedule information is as shown in Figure 8 In FIG. (b) below, as shown by 70, the schedule information 70 includes a subject, a train number, a departure time point, an arrival time point, a departure date, a location, etc.
[0073] It should be noted that the above application scenarios are all exemplary and do not limit the application scenarios of the method provided by the embodiments of the present application. In another embodiment, the mobile phone can also extract schedule information from other types of pictures, and other types of pictures include pictures other than instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots. Of course, other types of pictures can be pictures obtained by taking screenshots or pictures taken by a camera.
[0074] As can be seen from the above records, schedule information may come from different data sources, and the business scenarios involved include various types, such as IM chats, orders, IM meetings, voice assistants, etc. Different business scenarios result in different data formats. As an example, different system prompts can be set for different business scenarios, and the system prompts corresponding to each business scenario can include scenario description information, which is used to describe the business scenario. In this way, targeted instruction prompts are injected into the LLM through different system prompts, thereby improving the LLM's ability to extract schedule information.
[0075] Examples of some system prompts in the business of creating schedule information include:
[0076] (1) "IM Chat": "The following content is a chat conversation among several people, which may contain several events. Each event contains fields such as <'start time', 'end time', 'title','repeat cycle', 'location', 'url link', 'participants'>. The existing fields of each event are output in one line in JSON format without additional replies.\n\n"
[0077] (2) "IM Meeting": "The following content is some meeting notifications, which may contain several events. Each event contains fields such as <'start time', 'end time', 'title','repeat cycle', 'location', 'url link', 'participants'>. The existing fields of each event are output in one line in JSON format without additional replies.\n\n"
[0078] (3) "IM Notification": "The following content may contain several events. Each event contains fields such as <'start time', 'end time', 'title','repeat cycle', 'location', 'url link', 'participants'>. The existing fields of each event are output in one line in JSON format without additional replies.\n\n"
[0079] (4) "Voice Assistant": "The following is the content sent by the user to the voice assistant, which may contain several events. Each event contains fields such as <'start time', 'end time', 'title','repeat cycle', 'location', 'url link', 'participants'>. The existing fields of each event are output in one line in JSON format without additional replies.\n\n"
[0080] "(5) \"Other order screenshots\": \"The following is the content of the order page, which may contain several events. Each event contains fields such as <'start time','end time', 'title','repeat period', 'location', 'url link', 'participants'>. The fields that exist for each event are output in one line in json format, without additional response.\n\n"
[0081] The business scenarios corresponding to each system prompt in the above (1)-(5) are IM chat service, IM meeting service, IM notification service, voice assistant service, and order screenshot service in sequence.
[0082] As an example of this application, please refer to Figure 9 , the electronic device can calculate the KV value corresponding to each system prompt in multiple pre-set system prompts offline. Each system prompt in the multiple system prompts corresponds to a business scenario. For example, the electronic device can input each system prompt in the multiple system prompts into a pre-trained LLM for attention calculation to determine the KV value corresponding to each system prompt. Then, the electronic device caches the KV value corresponding to each system prompt. Exemplarily, the KV value corresponding to each system prompt is cached in the form of a matrix, that is, the matrix of KV values between characters in each system prompt is cached. In this way, during the application process, the electronic device no longer inputs the system prompt into the LLM. That is, the input is not the system prompt and the user prompt, but reads the KV value corresponding to the required system prompt from the cached KV values as a part of the model input, that is, the input is the user prompt and the read KV value. In this way, the LLM does not need to calculate the KV value of the system prompt in real time, reducing the content that the LLM needs to calculate in real time during application, thereby improving the inference speed of the LLM.
[0083] Please continue to refer to Figure 9 , taking the object to be extracted for schedule information as the first picture as an example, during the application process, it mainly includes the following steps 1-7:
[0084] 1. Obtain the first picture to be extracted for schedule information.
[0085] The first picture can be obtained by the electronic device taking a screenshot of the application interface of a certain application, such as taking a screenshot of the chat interface of the WeChat application; or the first picture can also be obtained by being sent from another electronic device.
[0086] 2. Receive the first picture dragged by the user through the calendar application.
[0087] Exemplarily, refer toFigure 4 In an embodiment, the user can submit the first picture to the calendar application by dragging.
[0088] 3. Determine the business scenario corresponding to the first picture.
[0089] As an example of this application, the electronic device can determine the business scenario corresponding to the first picture according to one or more of the business entry information, application information, and intent description information related to the first picture.
[0090] The business entry information is used to indicate the entry used to request text information inference. Exemplarily, when the user drags the first picture into the calendar application to request schedule extraction, the calendar application information is the business entry information related to the first picture.
[0091] The application information can be used to indicate the category of the application from which the first picture comes. The application information can be the name or category of the application, etc. Exemplarily, in a possible scenario, refer to Figure 2 , the first picture is obtained by the electronic device taking a screenshot of the chat interface in the foreground running instant messaging application. In this case, it can be determined that the application information related to the first picture is the application information of the instant messaging application.
[0092] The intent description information is used to indicate the intent of the first picture. As an example, the electronic device can determine the intent description information corresponding to the first picture through a pre-trained intent classification model. The pre-trained intent classification model can be obtained by training a neural network model based on training data, and the embodiments of this application do not limit this.
[0093] As an example of this application, the specific implementation of determining the business scenario corresponding to the first picture may include: in the case where the first picture corresponding to the target text information requests schedule extraction through the calendar application, if the first picture is obtained by the electronic device taking a screenshot of the foreground application, then according to the application information of the foreground application, determine the business scenario corresponding to the target text information. If the first picture is not obtained by the electronic device taking a screenshot, then determine the intent description information corresponding to the first picture through a pre-trained intent classification model, and according to the intent description information corresponding to the first picture, determine the business scenario corresponding to the target text information.
[0094] That is, in the case where the first picture requests schedule extraction through the calendar application, the corresponding business scenario can be first judged according to the application information. If it cannot be determined according to the application information, then determine the corresponding business scenario by determining the intent description information and according to the intent description information.
[0095] Exemplarily, refer to Figure 4, the first picture is obtained by the electronic device taking a screenshot of the chat interface in the instant messaging application running in the foreground. The user drags the first picture to the calendar application to trigger the electronic device to perform schedule extraction. In this case, the electronic device can know that a screenshot operation has been performed on the instant messaging application, and requests schedule extraction through the calendar application, so that the business entry information related to the first picture can be determined as the calendar application information, and the application information related to the first picture can be determined as "IM chat category", and then the business scenario corresponding to the first picture can be determined as "IM chat business".
[0096] Exemplarily, the first picture may also be obtained by other electronic devices through screenshot and then sent to this electronic device, for example, sent to this electronic device through the instant messaging application, and then the electronic device requests schedule extraction through the calendar application. In this case, the electronic device cannot know which application the first picture is a screenshot of by other electronic devices, that is, it cannot determine the application information corresponding to the first picture. At this time, the electronic device can determine the intention description information of the first picture through a pre-trained intention classification model. For example, if the intention description information indicates that the intention of the first picture is IM chat, the business entry information related to the first picture can be determined as the calendar application information, and the intention description information related to the first picture can be determined as IM chat, so that the corresponding business scenario can be determined as "IM chat business". Another example is that if the intention description information indicates that the intention of the first picture is an order, the corresponding business scenario can be determined as "order screenshot business".
[0097] It should be noted that the above is described by taking the object to be scheduled information extraction as a picture as an example. In another example, if the object to be scheduled information extraction is text information, for example, text information dragged from text, the electronic device can determine the corresponding business scenario according to the keyword or business entry information related to the text information.
[0098] Among them, different business scenarios can correspond to different keywords, and the keywords can be set in advance according to needs. For example, the keywords include voice assistant, IM chat, order, etc. Exemplarily, in a possible scenario, when the object to be scheduled information extraction is text information, if the text information includes the keyword "create meeting schedule", the electronic device can determine the business scenario corresponding to the text information as "IM meeting business" according to this keyword.
[0099] In a possible scenario, the text information is submitted through the voice assistant. That is, the voice assistant receives the voice, converts the voice into text information, and then requests the electronic device to extract the schedule from the text information. At this time, the business entry information related to the text information obtained can be "voice assistant", and according to this business entry information, the business scenario corresponding to the text information can be determined as "voice assistant service".
[0100] It should be noted that when the object for which the schedule information is to be extracted is the first picture, it is necessary to determine the target text information corresponding to the first picture, and then the schedule information is extracted based on the target text information through the LLM. Therefore, the business scenario corresponding to the first picture can also be referred to as the business scenario corresponding to the target text information. In the embodiments of the present application, the electronic device can select, through a certain strategy, which of the keyword, business entry information, application program information, and intent description information to use to determine the business scenario corresponding to the text information. Exemplarily, it can be determined which item to use for judgment according to the source of the text information. For example, if the text information is from a picture, in this case, the business scenario can be determined according to the application program information and business entry information related to the picture. If the application program information cannot be determined, the business scenario is determined through the intent description information. If the text information is from voice, such as a schedule extraction request through the voice assistant, the business scenario can be directly determined according to the business entry information. Another example is that if the text information is from text, the business scenario can be determined according to the keyword in the text information.
[0101] 4. Determine the corresponding system prompt according to the business scenario of the first picture.
[0102] As described above, different business scenarios correspond to different system prompts. After determining the business scenario corresponding to the first picture, the corresponding system prompt can be determined from multiple system prompts. The system prompt corresponding to the first picture is the system prompt corresponding to the target text information.
[0103] 5. Obtain the KV value corresponding to the determined system prompt from the cache.
[0104] As described above, the electronic device caches the KV values corresponding to different system prompts. After determining the system prompt corresponding to the target text information, the KV value corresponding to the determined system prompt can be obtained from the cache.
[0105] As an example of this application, the electronic device can perform the operations in 3-5 above through the service distribution engine. The service distribution engine can respond quickly after being triggered and is used to obtain the corresponding KV value according to the service scenario. In this way, by distributing the task of determining the KV value to the service distribution engine for execution, the electronic device can perform other operations in parallel, such as performing the operations in step 6 below, thereby improving the data processing efficiency of the electronic device.
[0106] 6. Obtain target text information based on the first picture.
[0107] The electronic device obtains the text information in the first picture to be extracted for schedule information, and obtains the target text information.
[0108] In implementation, the electronic device can perform content recognition and format parsing on the first picture, and then perform preprocessing to obtain the target text information. The preprocessing includes processing such as format, layout, line breaking, and wrapping. The specific implementation can be seen in the embodiments shown below. Figure 12 as shown in the embodiments.
[0109] 7. Use the target text information and the obtained KV value as the input of the LLM, and extract schedule information through the LLM.
[0110] In one example, the electronic device inputs the obtained KV value and the target text information into the LLM to extract schedule information through the LLM. That is, in the embodiments of this application, the input data of the LLM does not include the system prompt, but is replaced by the KV value corresponding to the system prompt.
[0111] In this way, by changing the model input from the original user prompt + system prompt to the KV value of the user prompt + system prompt, the computational amount of the LLM can be reduced. For example, if the data length of the user prompt + system prompt is 200 tokens, where the user prompt is 100 tokens and the system prompt is 100 tokens, the LLM only needs to calculate the KV value of the 100 tokens of the user prompt and the input KV value, which can save the computational amount of the 100 tokens of the system prompt and reduce the latency of calculating 100 tokens. For example, if the time for calculating a single token is K, the latency of K * L seconds can be reduced, where K is the length of the system prompt, so that the LLM can quickly extract schedule information.
[0112] It should be noted that the above several steps are only exemplary. In application, there may also be other operations, such as operations including extraction and preprocessing of target text information, display of schedule information, etc. Exemplarily, see Figure 10 ,Figure 10 It is a schematic diagram of a method for displaying schedule information shown according to an exemplary embodiment. In implementation, the electronic device can pass the first picture to be processed to the calendar application through interaction entrances such as sharing, Any Door, global collection, or Magic Text. The first picture can be a screenshot of an order for a train ticket, airplane ticket, or a screenshot of an order for a hotel, catering, entertainment, or a screenshot of an order for sports, health, or a screenshot of a WeChat mini-program, web, etc., or a screenshot of a service notification card of an IM notification number, or a screenshot of an IM chat conversation, etc. By way of example and not limitation, the screenshot of the order for a train ticket or airplane ticket can be from applications such as 12306, China Railway, or Ctrip, and the screenshot of the order for a hotel, catering, entertainment can be from applications such as Ctrip, Qunar, Tongcheng, Meituan, Fliggy, Dianping, Damai, etc., and the screenshot of the order for sports, health can be from applications such as Keep, Lianduoduo, registration platforms, etc., and the screenshot of a WeChat mini-program, web, etc. can include content such as performances, flash sales, marathons, etc., and the screenshot of the service notification card of an IM notification number can include notification cards issued by third parties such as hospitals, scenic spot tickets, educational institutions, insurance companies, etc. through the IM notification number, and the screenshot of an IM chat conversation can include content such as work arrangements, invitations, educational tasks, etc.
[0113] After the calendar application receives the first picture, it requests text recognition and edge detection of the first picture. Through text recognition, the text information in the first picture can be extracted, and through edge detection, the chat elements or color blocks in the picture can be recognized. Refer to Figure 10 , after passing the picture p1 to the calendar application through the interaction entrance, and after the calendar application requests the electronic device to perform text recognition, the obtained text recognition result is as shown in Figure 10 90 in.
[0114] After that, the electronic device can perform filtering processing on the text recognition result and the edge recognition result based on the filtering rules, that is, perform pre-processing to remove the interference information irrelevant to the schedule in the text recognition result and the edge recognition result. As an example of the application, the preset filtering rules can include at least one of the following rules: 1. Discard dense text blocks. 2. Discard small characters. 3. Discard skewed lines. 4. Discard floating text on the drawing. 5. Discard the "back" in the upper left corner of the picture and the "+" symbol in the picture.
[0115] It should be noted that in the process of extracting the schedule from the first picture, if the text recognition result is not filtered but directly input into the LLM for schedule information extraction, the recognition success rate of the LLM is relatively low, usually only reaching about 30%. The reason is the problems of interference information, format line breaks, and loss of layout information. Therefore, in the embodiments of the present application, filtering the text recognition result before extracting the schedule information through the LLM can improve the accuracy of the subsequent LLM in extracting the schedule information.
[0116] In addition, after receiving the first picture, the calendar application also requests to determine the business scenario corresponding to the first picture, and determines the corresponding system prompt according to the business scenario, so as to obtain the corresponding KV value according to the determined system prompt.
[0117] After the filtering process, the electronic device splices the target text information based on the remaining text recognition result and the edge recognition result after filtering, and uses the obtained KV value and the target text information as inputs and inputs them into the LLM to extract schedule information through the LLM. The LLM outputs schedule information based on the input prompt. Exemplarily, the output result is as shown in Figure 10 91 in. After that, the electronic device performs post-processing on the schedule fields of the output schedule information to obtain the schedule information to be displayed. Exemplarily, the schedule information to be displayed is as shown in Figure 10 92 in. After that, the electronic device can display the schedule information.
[0118] In one example, the post-processing of the schedule fields may include but is not limited to at least one of the following: 1. Processing according to the reminder time rule. 2. Processing according to the start time rule of tomorrow. 3. Processing according to the details rule. 4. Processing based on the pre-filling rule of title merging.
[0119] The reminder time rule includes setting the reminder time of the schedule information earlier than the time information extracted by the LLM by a preset duration. The preset duration can be set according to requirements. For example, the preset duration is 30 minutes. If the start time point extracted by the LLM is 8:30, the reminder time of the schedule information can be set to 8:00. In addition, the reminder time rule also includes repeated reminders. For example, if the schedule information extracted by the LLM is to grab numbers on Monday, Tuesday, and Wednesday, the electronic device sets the reminder time for Monday, the reminder time for Tuesday, and the reminder time for Wednesday for this schedule information, rather than setting only one reminder time.
[0120] The tomorrow start time rule means that if there is a situation of crossing days, months or years in the time information extracted by the LLM, the time after crossing days, months or years is supplemented. For example, if the date extracted by the LLM is December 5th, the start time point is 23:00, and the end time point is 00:30, the electronic device can supplement the end time point as 00:30 on December 6th in the schedule information.
[0121] The details rule means that the layout and font size of the created schedule information are adjusted according to the size of the screen of the electronic device, so that it can be correctly and clearly displayed in the schedule details area of the calendar application.
[0122] The title merging pre-filling rule includes, in the case where there are multiple different titles corresponding to the same time information, selecting the title with the longest length as the title of the schedule information from the multiple titles. In addition, the title merging pre-filling rule also includes using the specified title corresponding to the scenario of using a picture as the title of the schedule information. For example, if the scenario of the picture is making an appointment to get a number for seeing a doctor, and the title extracted by the LLM is "Getting a number" or "Taking a number", the title of the schedule can be standardized as "Registering for a medical appointment", where the specified titles corresponding to different scenarios can be preset according to requirements.
[0123] It should be noted that the above post-processing of the schedule fields is only exemplary. In another example, the post-processing of the schedule fields may also include, but is not limited to, at least one of time similarity processing, discarding empty results, risk control, date standardization, time standardization, symbol standardization, and cleaning non-natural language titles. Time similarity processing means that if the time in the output schedule information is earlier than the current system time of the electronic device, the time closest to the time in the schedule information is determined according to the current system time, and the determined time is determined as the time in the schedule information. For example, if the time in the schedule information is Tuesday and the current system time is Wednesday, it is recorded as Tuesday of the next week in the schedule information. Discarding empty results means discarding the returned empty fields. Risk control means controlling sensitive words. Date standardization means expressing the date in the form of xx year xx month xx day. Time standardization means adding information such as am (morning) or pm (afternoon) to the time point. Symbol standardization means unifying the symbols into the same style. Cleaning non-natural language titles means that if the title of the schedule does not include a verb, a verb can be added to the title of the schedule or the title of the schedule can be set by default according to the scenario corresponding to the picture.
[0124] The software system of the electronic device involved in the embodiments of the present application can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. Taking the Android system with a layered architecture as an example, the software system of the electronic device is exemplarily described in the embodiments of the present application.
[0125] Figure 11 It is a block diagram of a software system of an electronic device provided by an embodiment of the present application. Refer to Figure 11 , the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The communication between layers is through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system layer, and the kernel layer.
[0126] The application layer may include a series of application packages. Refer to Figure 11 , the application packages may include a calendar and other application programs. For example, other application programs may include instant messaging, ticket booking, camera, gallery, call, map, navigation, Bluetooth, music, video, short message, voice assistant and other application programs.
[0127] In addition, as an example of the present application, the application layer further includes a schedule management service and a model management service. The schedule management service can be used to provide services for the calendar, or be called by the calendar. For example, the schedule management service can create schedule information for the calendar and store data such as schedule information. The schedule management service may include a schedule creation service (which can be called: intelligent parsing and processing service) and a schedule database (such as: Calednar Provider schedule database). The schedule management service can create schedule information through the schedule creation service and store the created schedule information through the schedule database. The model management service (which can be called: MagicLive large model service) can be used to provide various models for the schedule management service to call when needed.
[0128] In one example, the model management service provides a pre-trained LLM, an object recognition model, and a personal behavior feature model. Among them, the LLM can be used to recognize (i.e., reason about) the prompt to determine the schedule information. Exemplarily, the LLM can be a natural language understanding (NLP) model, and the NLP model can run through a natural language unit (NLU). In the embodiments of the present application, the pre-trained LLM is referred to as the target natural language model. In some examples, the target natural language model can not only recognize the prompt but also extract keywords in a piece of text, such as keywords like time and location. The object recognition model can be used for text recognition and edge recognition of pictures, and can also be used to determine the picture category of the pictures. In one example, the object recognition model includes a first optical character recognition (OCR) model, a second OCR model, and an edge detection model. The first OCR model can be used for text recognition, the second OCR model can be used to determine the picture category, and the edge detection model can be used for edge recognition of pictures. The number of edge detection models can be multiple, and different edge detection models can be used for edge recognition of pictures of different picture categories; the personal behavior feature model can be used to determine the user portrait according to the user's historical behavior data.
[0129] It should be noted that the embodiments of the present application are described by taking the target natural language model as a pre-trained LLM as an example. In another example, the size of the model and whether it is multimodal may not be limited.
[0130] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. As Figure 11 shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0131] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core libraries consist of two parts: one is the functional functions that the Java language needs to call, and the other is the core libraries of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0132] The system libraries can include multiple functional modules, such as: Surface Manager, Media Libraries, 3D graphics processing libraries (such as OpenGL ES), 2D graphics engines (such as SGL), etc.
[0133] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, audio drivers, and sensor drivers.
[0134] The electronic device can implement the method for displaying schedule information provided in the embodiments of the present application through the interaction of the above multiple modules. Next, in combination with Figure 12 The process of extracting and displaying schedule information in the picture will be introduced in detail. See Figure 12 This method may include the following implementation steps:
[0135] S1201: The calendar application receives the picture L to be processed.
[0136] The picture L can be a picture in bitmap format.
[0137] The calendar application receives the picture L passed by the interaction entry. As mentioned above, the interaction entry can be a sharing entry, an arbitrary door entry, etc. Exemplarily, see Figure 2 In the figure (a) of, when the user submits the picture L to the calendar application through the sharing control, this interaction entry is the sharing entry.
[0138] S1202: The calendar application sends a schedule creation instruction to the schedule management service, and the picture L is carried in the schedule creation instruction.
[0139] The schedule creation instruction is used to indicate the creation and display of relevant schedule information based on the picture L.
[0140] S1203: The schedule management service determines the business scenario corresponding to the picture L.
[0141] S1204: The schedule management service determines the corresponding target system prompt according to the business scenario corresponding to the picture L.
[0142] S1205: The schedule management service obtains the KV value corresponding to the target system prompt from the cached KV values.
[0143] The cached KV values include the KV value corresponding to each system prompt among multiple system prompts.
[0144] As an example of the present application, the schedule management service can trigger the business distribution engine to execute the operations of S1203 to S1205, and the specific implementation can be referred to Figure 9 the embodiments shown.
[0145] S1206: The schedule management service sends picture L to the first OCR model in the model management service.
[0146] In implementation, the schedule management service invokes the first OCR model in the model management service and sends picture L to the first OCR model for text recognition processing.
[0147] S1207: The first OCR model determines the first recognition result of picture L.
[0148] As an example of the present application, the first recognition result includes the text recognition result of picture L, and the text recognition result includes text block coordinates, text line coordinates, and text line recognition content.
[0149] As an example, the first recognition result further includes target indication information, and the target indication information can be used to indicate whether the picture L input into the first OCR model is a screenshot picture or a taken picture. Exemplarily, the target indication information can be a first identifier, a second identifier, or a third identifier. The first identifier is used to indicate that the picture L input into the first OCR model is a screenshot picture, the second identifier is used to indicate that the picture L input into the first OCR model is a taken picture and is a picture taken of a document, and the third identifier is used to indicate that the picture L input into the first OCR model is other taken pictures, such as pictures taken of advertisements, road signs, magazines, etc. The first identifier, the second identifier, and the third identifier can be set according to requirements. For example, the first identifier is F1, the second identifier is F2, and the third identifier is F3.
[0150] That is, after receiving picture L, the first OCR model recognizes picture L and outputs the first recognition result of picture L.
[0151] S1208: The first OCR model sends the first recognition result to the schedule management service.
[0152] As an example, after receiving the first recognition result, the schedule management service can cache the first recognition result.
[0153] S1209: The schedule management service sends picture L to the second OCR model in the model management service.
[0154] In one example, after receiving a schedule creation instruction sent from a calendar application, in addition to sending picture L to the first OCR model for text recognition processing, the schedule management service can also call the second OCR model in the model management service and send picture L to the second OCR model to determine the picture category through the second OCR model. That is, the operations of S1206 and S1203 can be executed in parallel.
[0155] Since the method provided in the embodiments of the present application can extract schedule information from pictures of different picture categories, and the content layouts of pictures of different picture categories are different, the electronic device processes pictures of different picture categories in different ways. Therefore, in implementation, after receiving picture L to be processed, the schedule management service not only performs text recognition through the first OCR model, but also inputs picture L into the second OCR model to determine the picture category of picture L.
[0156] As an example of the present application, the picture categories include instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. An instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application; an instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application; an order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application; other category pictures include other pictures except instant messaging chat screenshots, instant messaging notification card screenshots, and order screenshots, and other category pictures can be screenshot pictures or captured pictures.
[0157] S1210: The second OCR model determines the picture category of picture L.
[0158] After receiving picture L, the second OCR model recognizes picture L and outputs the picture category of picture L.
[0159] S1211: The second OCR model sends the picture category of picture L to the schedule management service.
[0160] S1212: The schedule management service sends picture L to the edge detection model corresponding to the picture category of picture L.
[0161] As an example of this application, multiple edge detection models are provided in the model management service. The multiple edge detection models can be pre-trained, and different edge detection models can identify the edges of pictures of different picture categories. In one example, the multiple edge detection models include a first edge detection model and a second edge detection model. The first edge detection model can be used to identify the edges of instant messaging chat screenshots to determine the coordinates and categories of chat elements in the instant messaging chat screenshots; the second edge detection model can be used to identify the edges of pictures other than instant messaging chat screenshots. For example, the second edge detection model can be used to identify the edges of instant messaging notification card screenshots, order screenshots or other category pictures to determine the coordinates and categories of color blocks in other pictures.
[0162] The different edge detection models can be obtained by iteratively training an initial training model based on picture training samples of pictures of the corresponding picture categories. The picture training samples can be obtained by edge annotation in advance according to requirements. The initial training model can be set according to requirements. As an example but not a limitation, the initial training model can be a recurrent neural network (RNN), etc.
[0163] After the schedule management service receives the picture category of picture L, it determines the edge detection model corresponding to the picture category of picture L from the multiple edge detection models. Exemplarily, when picture L is p1 in Figure (a) of Figure 3 , that is, an instant messaging chat screenshot, the edge detection model corresponding to the picture category of picture L is determined from the multiple edge detection models as the first edge detection model; when picture L is p3 in Figure (a) of Figure 8 (which is an instant messaging notification card screenshot) or p4 in Figure (a) of Figure 9 (which is an order screenshot), the edge detection model corresponding to the picture category of picture L is determined from the multiple edge detection models as the second edge detection model. Then, the schedule management service calls the determined edge detection model and sends picture L to the edge detection model to request edge recognition of picture L.
[0164] S1213: The edge detection model performs edge recognition processing on picture L and outputs an edge recognition result.
[0165] The edge recognition result includes graphic and text element attribute information or color block attribute information.
[0166] In one example, when the picture category of picture L is an instant messaging chat screenshot, the edge detection model corresponding to the picture category of picture L is the first edge detection model. After inputting picture L into the first edge detection model for processing, the output edge recognition result includes graphic and text element attribute information, and the graphic and text element attribute information includes the attribute information of the chat elements in the instant messaging chat screenshot. The attribute information of the chat elements may include the coordinates and categories of the chat elements. Exemplarily, the chat elements include avatars, titles, nicknames, chat content, chat timestamps, usernames, specified identifiers, etc., and the specified identifier includes the "+" identifier. For example, see Figure 13 , after inputting picture L into the first edge detection model, the first edge detection model can determine that the chat elements in picture L include Figure 13 the multiple items identified by the dashed box in
[0167] . Each chat element corresponds to its own coordinates and category. For example, the coordinates of a certain chat element are the coordinates of the four corners of the area where the chat element is located, and the category is an avatar. Optionally, the attribute information of each chat element may further include the serial number of the chat element, and the serial numbers of the chat elements in picture L can be set by default according to the sorting of the chat elements in picture L. Figure 14 , when picture L is an instant messaging notification card screenshot, after inputting picture L into the second edge detection model, the second edge detection model can determine that the color blocks in picture L include Figure 14 the multiple items identified by the dashed box in
[0168] S1214: The edge detection model sends the edge recognition result to the schedule management service.
[0169] As an example, after receiving the edge recognition result, the schedule management service can cache the edge recognition result.
[0170] It is worth mentioning that after performing text recognition and edge recognition processing on picture L through the above two branches respectively, a first recognition result and a second recognition result can be obtained. The first recognition result includes a text recognition result and target indication information, and the second recognition result includes an edge recognition result and a picture category. Since the text recognition result can represent the text content in picture L and the edge recognition result can represent the layout of picture L, subsequent extraction of schedule information based on these two types of data, namely the first recognition result and the second recognition result, can improve the accuracy of information extraction. The specific implementation can refer to the following steps.
[0171] It should be noted that there is no strict order of execution among the operations of S1203 to S1205, the operations of S1206 to S1207, and the operations of S1208 to S1214. In one example, the three branches can be executed in parallel.
[0172] S1215: When the picture category of picture L is an instant messaging chat screenshot, the schedule management service filters the text recognition result and the edge recognition result according to the first filtering rule.
[0173] Different picture categories correspond to different filtering rules. In practice, the schedule management service determines the corresponding filtering rule according to the picture category of picture L, and then filters the text recognition result and the edge recognition result of picture L according to the determined filtering rule to filter out interference information irrelevant to the schedule.
[0174] In one example, the instant messaging chat screenshot corresponds to the first filtering rule. The first filtering rule may include filtering out the recognition data corresponding to the skewed lines, small characters, and chat timestamps respectively. A skewed line refers to a text line corresponding to a chat element with an inclination angle greater than a preset angle, and the preset angle can be set according to requirements. For example, the preset angle is 10 degrees; small characters refer to a text line corresponding to a chat element with a line height less than the target line height, and the target line height may refer to the average line height of all text lines in picture L; the chat timestamp is a timestamp used to indicate the chat time. Optionally, the first filtering rule may also include filtering out dense text blocks, floating text on the picture, "back" in the upper left corner, "+", etc. in the lower right corner. Exemplarily, see Figure 13 , Figure 13 According to an exemplary embodiment, some chat elements in picture L that need to be filtered are identified, including chat timestamp 1101, small characters 1102, skewed line 1103, and "+" 1004.
[0175] As an example, in the implementation of filtering out the recognition data corresponding to skewed lines, the inclination angle of each text line can be calculated based on the text line coordinates in the text recognition result, so as to determine which text lines are skewed lines, and then delete the recognition data corresponding to the skewed lines. For example, delete the text line coordinates and text line recognition content of the skewed lines.
[0176] As an example, in the implementation of filtering out the recognition data corresponding to small characters, the line height of each text line in the picture L can be determined according to the text line coordinates in the text recognition result, so as to filter out the recognition data corresponding to the text lines with a line height less than the target line height. For example, filter out the text line coordinates and text line recognition content of the text lines with a line height less than the target line height.
[0177] As an example, in the implementation of filtering out the recognition data corresponding to chat timestamps, chat elements whose category is chat timestamps can be determined according to the edge recognition result, and at least one first candidate chat element can be obtained. According to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the picture L, match the text line recognition content corresponding to each first candidate chat element in the text recognition result of the picture L. When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element. The target first candidate chat element is any one of the at least one first candidate chat elements. When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filter out the recognition data corresponding to the target first candidate chat element from the text recognition result and the edge recognition result of the picture L.
[0178] As described above, the edge recognition result of the instant messaging chat screenshot includes the category of chat elements. Therefore, according to the category of chat elements in the edge recognition result, chat elements with the category of chat timestamp can be filtered out to obtain at least one first candidate chat element. Since the edge recognition result does not include the text line recognition content, that is, it is impossible to know the text content corresponding to each chat element. In some possible cases, the edge detection model may misjudge the category of chat elements that are not chat timestamps as chat timestamps. Therefore, in order to avoid filtering out chat elements that are not chat timestamps as much as possible, after determining at least one first candidate chat element, the text line recognition content corresponding to each first candidate chat element can be matched from the text recognition result of picture L according to the coordinates of each first candidate chat element and the text line coordinates of picture L. For example, for any one of the first candidate chat elements, according to the coordinates of this first candidate chat element and the text line coordinates of picture L, the text line recognition content of at least one text line located in the area corresponding to this first candidate chat element is determined from the text recognition result of picture L, so as to match the text line recognition content corresponding to this first candidate chat element. After that, the schedule management service can determine whether the matched text line recognition content is a timestamp through the target natural language model. For example, the schedule management service can send the matched text line recognition content to the target natural language model through getEntitiy() to request the target natural language model to identify whether the text line recognition content is a timestamp. If it is determined through the target natural language model that the text line recognition content is a timestamp, it can be further determined that this first candidate chat element may be a chat timestamp. Otherwise, if it is determined through the target natural language model that the text line recognition content is not a timestamp, it can be determined that this first candidate chat element is not a chat timestamp.
[0179] Since the position of the chat timestamp in the instant messaging chat screenshot is generally fixed. For example, it is usually centered in the chat window of WeChat. Therefore, in the case where it is determined through the target natural language model that a certain first candidate chat element may be a chat timestamp, it can be further determined whether this first candidate chat element is a chat timestamp according to the coordinates of this first candidate chat element, so as to improve the accuracy of chat timestamp recognition.
[0180] As an example of the present application, for any one of the at least one first candidate chat elements, the specific implementation of determining whether to filter out the recognition data corresponding to this first candidate chat element according to the coordinates of this first candidate chat element may include the following two cases:
[0181] The first case: When it is determined according to the coordinates of this first candidate chat element that this first candidate chat element is located in the middle position of Picture L and the area corresponding to this first candidate chat element includes a single text line, determine to filter out the recognition data corresponding to this first candidate chat element.
[0182] Since in the chat windows of most instant messaging applications, the chat timestamp is generally centered and the text line recognition content in the chat timestamp only includes one line, that is, only the timestamp. Therefore, if it is determined that this first candidate chat element is located in the middle position of Picture L and the area corresponding to this first candidate chat element only includes a single text line, and since it has been determined through the target natural language model that the text line recognition content is the timestamp, it can be determined that this first candidate chat element is the chat timestamp, that is, it is determined that the recognition data corresponding to this first candidate chat element can be filtered out.
[0183] The second case: When it is determined according to the coordinates of this first candidate chat element that this first candidate chat element is located on the right side of Picture L and the area corresponding to this first candidate chat element includes a single text line, if the text line recognition content corresponding to this first candidate chat element only includes time and does not include date, determine to filter out the recognition data corresponding to this first candidate chat element.
[0184] Since in the chat windows of some instant messaging applications (such as the chat interface forwarded in WeChat), the chat timestamp may also be displayed on the right side, and the text line recognition content in the chat timestamp only includes one line. Additionally, the chat timestamp only includes time and does not include date. Therefore, if it is determined that this first candidate chat element is located on the right side of Picture L and the area corresponding to this first candidate chat element only includes a single text line, it can be determined whether the text line recognition content corresponding to this first candidate chat element includes a date. For example, it can be determined whether the text line recognition content includes a date through the target natural language model. If it is determined that it does not include a date, it can be determined that this first candidate chat element is the chat timestamp, that is, it is determined that the recognition data corresponding to this first candidate chat element can be filtered out. Of course, if it is determined that it includes a date, it can be determined not to filter.
[0185] In one example, for the second case, it is also possible not to determine whether the text line recognition content corresponding to this first candidate chat element only includes time and does not include date. As long as it is determined that this first candidate chat element is located on the right side of Picture L and the area corresponding to this first candidate chat element includes a single text line, the schedule management service can determine to filter out the recognition data corresponding to this first candidate chat element.
[0186] When it is determined through the above process that a certain first candidate chat element is a chat timestamp, the schedule management service deletes the recognition data corresponding to this first candidate chat element from the text recognition result and the edge recognition result of picture L. For example, it deletes the text line coordinates and the recognized text content corresponding to this first candidate chat element from the text recognition result of picture L, and deletes the coordinates and category corresponding to this first candidate chat element from the edge recognition result of picture L. Of course, if it is determined through the above process that a certain first candidate chat element is not a chat timestamp, the schedule management service does not filter out the recognition data corresponding to this first candidate chat element.
[0187] It is worth mentioning that first determining the chat elements whose category is a chat timestamp according to the edge recognition result, then matching the corresponding recognized text content from the text recognition result, determining whether it is a chat timestamp through the target natural language model based on the matched recognized text content, and then determining whether it is a chat timestamp according to the position of the chat element can improve the accuracy of chat timestamp recognition, thereby improving the accuracy of filtering, and further improving the accuracy of schedule information creation.
[0188] It should be noted that the above first filtering rule is only exemplary. When the instant messaging chat screenshots are from different instant messaging applications, the layout of their chat elements is usually different, and the chat elements may also be different, so that the interference information in the instant messaging chat screenshots of different instant messaging applications may be different. Exemplarily, it usually includes several possible situations shown in Table 3. Therefore, in another example, the first filtering rule may also include other rules for filtering out information irrelevant to the chat content.
[0189] Table 3
[0190]
[0191] In order to effectively filter out the interference information in the instant messaging chat screenshots from different instant messaging applications, a first filtering rule can be set according to the union of the possible interference information shown in Table 3, so as to ensure that no matter which instant messaging chat screenshot is processed, the interference information can be effectively removed. Exemplarily, the first filtering rule may further include filtering out the text in the avatar, the user name, the specified identifier, etc. In implementation, after filtering out the recognition data corresponding to the chat timestamp in the picture L from the text recognition result and the edge recognition result of the picture L, the schedule management service may determine, according to the categories of the remaining chat elements in the edge recognition result, the chat elements irrelevant to the chat content from the remaining chat elements in the edge recognition result, to obtain at least one second candidate chat element, and match the text line recognition content corresponding to each second candidate chat element in the text recognition result of the picture L according to the coordinates of each second candidate chat element in the at least one second candidate chat element and the text line coordinates in the text recognition result of the picture L. Delete the matched text line recognition content and the corresponding text line coordinates from the text recognition result of the picture L.
[0192] Exemplarily, in the implementation of filtering out the text in the avatar, the schedule management service may determine the chat element whose category is the avatar according to the edge recognition result, and then, according to the coordinates of the chat element, match the text line recognition content corresponding to the chat element in the text recognition result of the picture L. If there is a matched text line recognition content, delete the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content from the text line recognition result of the picture L, so as to delete the text in the avatar.
[0193] Exemplarily, in the implementation of filtering out the user name, the schedule management service may determine the chat element whose category is the user name according to the edge recognition result, and then, according to the coordinates of the chat element, match the text line recognition content corresponding to the chat element in the text recognition result of the picture L. If there is a matched text line recognition content, delete the matched text line recognition content and the text line coordinates corresponding to the matched text line recognition content from the text line recognition result of the picture L, so as to delete the user name.
[0194] Exemplarily, in the implementation of filtering out the specified identifier, the schedule management service may determine the chat element whose category is the specified identifier according to the edge recognition result, and then match the text line recognition content corresponding to the chat element in the edge recognition result of the picture L. If the text line recognition content is the specified identifier, for example, it is "+", the recognition data corresponding to the chat element may be deleted from the text line recognition result of the picture L, for example, the text line coordinates and the text line recognition content corresponding to the chat element are deleted. The schedule management service may also delete the recognition data corresponding to the chat element from the edge recognition result, for example, delete the coordinates and the category of the chat element.
[0195] It should be noted that the above is described by taking the picture L as a screenshot of an instant messaging chat as an example. In another example, if the picture L is not a screenshot of an instant messaging chat, for example, it is a screenshot of an instant messaging notification card or an order screenshot, the schedule management service filters out interference information based on the second filtering rule. In an example, in the implementation of filtering based on the second filtering rule, the schedule management service can match the text line recognition content in each color block from the text recognition result of the picture L according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the picture L. For any one of the color blocks, if it is determined according to the text line recognition content in any one of the color blocks that the content related to the schedule is not included in any one of the color blocks, for example, the time information and location are not included, the recognition data corresponding to any one of the color blocks is filtered out from the edge recognition result and the text recognition result of the picture L. In an example, it can be determined whether the color block includes time information and location through a target natural language model. For example, the text line recognition content in the color block can be sent to the target natural language model one by one to request the target natural language model to determine whether the time information and location are included.
[0196] In the case where the picture L is a screenshot of an instant messaging notification card, it can be seen from Table 3 that the possible interference information includes small characters, skewed lines, and middle timestamps. Therefore, in an example, before filtering based on the second filtering rule, the recognition data corresponding to the skewed lines, small characters, and middle timestamps in the text recognition result of the picture L and the edge recognition result can be filtered out first, and then filtered out based on the second filtering rule. The filtering of the skewed lines, small characters, and middle timestamps can refer to the first filtering rule.
[0197] As an example of the present application, before filtering, it can also be queried whether the number of color blocks in the picture L is less than the number threshold. If the number of color blocks in the picture L is less than the number threshold, it means that there are no a large number of color blocks in the picture L. In this case, the electronic device can usually process the picture L, so filtering processing can be performed according to the second filtering rule. If the number of color blocks in the picture L is greater than or equal to the number threshold, it means that the picture L includes a large number of color blocks. In this case, instead of performing filtering processing, a prompt message can be displayed to guide the user to re-take a screenshot of the picture L, for example, guiding the user to intercept a part of the area including the schedule information in the picture L through the electronic device. The number threshold can be set according to requirements. For example, the number threshold can be 10.
[0198] As an example of this application, before filtering, the schedule management service can also determine whether there are dense text blocks based on the text block coordinates in the text recognition result of picture L. If there are dense text blocks, a prompt message can be displayed in the calendar application. The prompt message is used to prompt the user that there are dense text blocks, so that the user can re-crop picture L according to the needs. If there are no dense text blocks, the schedule management service performs the filtering operation.
[0199] Or in another example, during the filtering process, if the schedule management service determines that there are dense text blocks based on the text block coordinates in the text recognition result of picture L, the text line coordinates and text line recognition content in these text blocks are deleted.
[0200] It is worth mentioning that if no filtering is performed, it is likely to cause the subsequent target natural language model to be unable to accurately extract schedule information. For example, for Figure 10 picture p1 in, the time may be extracted as the chat timestamp "20:37", and the address may be extracted as "No. 5, Danling Street, Ningyuan Road, Haidian District, Beijing", etc. In the embodiments of this application, after determining the first recognition result and the second recognition result, filtering out the interference information in picture L can improve the accuracy of subsequent schedule information extraction.
[0201] As an example rather than a limitation, when picture L is a captured picture, since the captured picture may be skewed and include a background, etc., no filtering may be performed, and the text recognition result is used for text splicing later.
[0202] In another example, when picture L is a captured picture, the text recognition result and edge recognition result of picture L can also be filtered according to the second filtering rule, and the embodiments of this application do not limit this.
[0203] S1216: The schedule management service matches the text line recognition content of each remaining chat element from the filtered text recognition result based on the filtered edge recognition result.
[0204] The schedule management service matches the text line recognition content corresponding to each filtered chat element from the filtered text recognition result according to the coordinates of each chat element in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.
[0205] When picture L is not an instant messaging chat screenshot, the schedule management service matches the text line recognition content in each filtered color block from the filtered text recognition result according to the coordinates of each color block in the filtered edge recognition result and the text line coordinates in the filtered text recognition result.
[0206] S1217: The schedule management service determines the scene corresponding to picture L based on the text line recognition content of each remaining chat element and the picture category of picture L.
[0207] The scene corresponding to picture L refers to the scene involved in the text content in picture L.
[0208] In the case where it is determined that picture L is an instant messaging chat screenshot based on the picture category of picture L, if it is further determined to be a chat conversation according to the text line recognition content of each filtered chat element, the scene corresponding to picture L can be determined as the IM chat scene.
[0209] In addition, in the case where the picture category of picture L is other pictures, such as an instant messaging notification card screenshot or an order screenshot, the schedule management service can also combine the text line recognition content of each matched color block to determine the scene corresponding to picture L, such as the service notification card scene or the high-speed rail travel order scene.
[0210] S1218: The schedule management service performs chat conversation splicing on the text line recognition content of each filtered chat element based on the scene corresponding to picture L to obtain the target text information.
[0211] In implementation, the schedule management service performs operations such as line breaks, carriage returns, and splicing on the text line recognition content of each filtered chat element according to the scene corresponding to picture L. Since chat content is usually rather casual, for example, a complete sentence may be sent in multiple messages. In the case where it is determined that the scene corresponding to picture L is the IM chat scene, the schedule management service can splice these multiple messages into one sentence. Therefore, performing chat conversation splicing based on the scene corresponding to picture L can make the spliced text closer to natural language text. By way of example and not limitation, the target text information after chat conversation splicing is consistent with the content displayed in the smearing interface.
[0212] In one example, the spliced chat conversation content includes the conversation type, title, nickname, conversation content, etc., and the nickname can be customized. For example, the chat conversation content can be spliced in the following format:
[0213] Conversation topic: IM chat
[0214] Title: xx group chat
[0215] Nickname A: xxx
[0216] Nickname B: xxx
[0217] Nickname A: xxxx .......
[0219] It should be noted that S1217 to S1218 are optional operations. In another example, the schedule management service can also splice chat conversations based on the text line recognition content of each filtered chat element according to the picture category of picture L.
[0220] In addition, when picture L is other pictures, such as a screenshot of an instant messaging notification card or an order screenshot, the schedule management service splices the text line recognition content in each filtered color block according to the scenario corresponding to picture L, and its implementation can refer to the splicing of chat conversations.
[0221] S1219: The schedule management service uses the target text information and the KV corresponding to the target system prompt as input data and sends the input data to the target natural language model.
[0222] As described above, in the stage of obtaining picture L, the schedule management service can determine the business scenario corresponding to picture L, that is, it can determine the business scenario to which the target text information belongs, and has matched the KV corresponding to the target system prompt for this business scenario from the cached KV values. The electronic device can use the KV corresponding to the prompt and the target text information as input data, call the target natural language model, and send the input data to the target natural language model to request the target natural language model to extract schedule information.
[0223] It should be noted that the embodiments of this application are described by taking the target natural language model being deployed in the electronic device as an example. In another example, the target natural language model can also be deployed in the cloud, and the cloud can provide an interface for the electronic device to call the target natural language model. Thus, when the target natural language model is needed, the schedule management service can call the target natural language model through the provided interface, and the embodiments of this application do not limit this.
[0224] S1220: The target natural language model determines schedule information based on the input data.
[0225] In one example, the schedule information output by the target natural language model is as Figure 10 shown in 91.
[0226] S1221: The target natural language model sends the schedule information to the schedule management service.
[0227] Exemplarily, the target natural language model can send the schedule information to the schedule management service in JSON format.
[0228] S1222: The schedule management service performs post-processing on the schedule fields of the schedule information.
[0229] The post-processing of the schedule fields of the schedule information can refer to the foregoing.
[0230] In one example, before creating the schedule information, the schedule management service may also call the personal behavior feature model to request a query of the user's historical behavior data. For example, the historical behavior data includes historical locations, etc. Correspondingly, the personal behavior feature model returns the historical behavior data. In this way, the schedule management service can predict the places the user may go based on the historical behavior data, and then combine the schedule information fed back by the target natural language model to create the final schedule information. For example, the predicted address information is added to the schedule information.
[0231] S1223: The schedule management service displays the processed schedule information in the calendar application.
[0232] Exemplarily, when the picture L is a screenshot of an instant messaging chat, the schedule information displayed on the electronic device is as shown at 13 in FIG. (d) below. Figure 2 as shown in FIG. (d) below.
[0233] In one example, before displaying the schedule information, a request confirmation notice may be displayed first. After receiving the confirmation display instruction triggered by the user based on the request confirmation notice, the schedule management service then displays the schedule information in the calendar application.
[0234] As an example of the present application, the electronic device also supports the user to edit the displayed schedule information. Exemplarily, it supports the user to modify the title of the schedule information. For example, refer to the embodiment shown below. Figure 3 During this process, when the user triggers the electronic device to display the target interface, text that can be smeared needs to be displayed in the target interface. For this purpose, the schedule management service may also request the target natural language model to perform word combination processing on the spliced text. For example, the two words "I" and "we" are combined into "we" to facilitate text display in the target interface according to the combined words, thereby facilitating the user to smear. In practice, the schedule management service can call the target natural language model through getWordSegment() to request the target natural language model to perform word combination processing. In addition, the schedule management service can also call the target natural language model through getWordSegment() to request the target natural language model to determine the theme of the spliced text. For example, the target natural language model is specified to extract the theme entity (such as meeting, dinner). In this way, the schedule management service can display the theme in the target interface.
[0235] In an embodiment of the present application, in response to a schedule extraction operation, target text information to be extracted is obtained. A system prompt corresponding to the business scenario to which the target text information belongs is determined from multiple system prompts, and a target system prompt is obtained. The target system prompt is used to describe the task of the time to be inferred in the target text information. The KV value of the target system prompt is obtained from the KV values of the cached multiple system prompts. The target text information and the KV value of the target system prompt are used as inputs to a pre-trained target natural language model, and inference is performed through the target natural language model to output schedule information corresponding to the target text information. In this way, it is not necessary for the user to manually input schedule information item by item in the calendar application, improving the efficiency of creating schedule information. Moreover, during the application process, since the electronic device can read the required KV value from the cached KV values as part of the model input, the target natural language model does not need to perform real-time calculation on the system prompt, reducing the content that the target natural language model needs to perform real-time calculation during application, thereby improving the inference speed of the target natural language model.
[0236] It should be noted that the embodiment of the present application is described by taking the extraction of schedule information as an example. In another example, the method provided by the embodiment of the present application can also be applied to other scenarios of extracting events, such as answering questions, etc.
[0237] The electronic device involved in the embodiment of the present application may be a mobile phone, a sports camera (GoPro), a digital camera, a tablet computer, a desktop type, a laptop, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, etc., and the embodiment of the present application does not limit this.
[0238] Figure 15 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Refer to Figure 15, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0239] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0240] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0241] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.
[0242] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0243] It can be understood that the interface connection relationships shown between the various modules in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0244] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, the wireless communication module 160, etc.
[0245] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.
[0246] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.
[0247] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is an integer greater than 1.
[0248] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0249] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0250] The internal memory 121 can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0251] The electronic device 100 can implement audio functions, such as music playback, recording, etc., through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.
[0252] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 together form a touch screen, also known as the "touch display screen". The touch sensor 180K is used to detect touch operations acting thereon or in its vicinity. The touch sensor 180K can transmit the detected touch operations to the application processor to determine the type of touch event. Visual outputs related to the touch operations can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a position different from that of the display screen 194.
[0253] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.
[0254] The above are the optional embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the technical scope disclosed in the present application shall be included within the protection scope of the present application.
Claims
1. A method for obtaining schedule information, characterized in that, Applied to an electronic device, the method includes: In response to a schedule extraction operation, obtaining target text information, where the target text information includes the schedule information to be extracted; Determining, from multiple system prompts, the system prompt corresponding to the target text information to obtain a target system prompt, where the system prompt is used to describe the task of the event to be inferred, and each system prompt in the multiple system prompts corresponds to a business scenario; Obtaining the KV value of the target system prompt from the cached KV values of the multiple system prompts; Using the target text information and the KV value of the target system prompt as the input of a pre-trained target natural language model, and performing inference through the target natural language model to output the schedule information corresponding to the target text information.
2. The method according to claim 1, wherein The determining, from multiple system prompts, the system prompt corresponding to the target text information includes: Determining the business scenario corresponding to the target text information according to one or more of the keyword, business entry information, application information, and intent description information related to the target text information, where different keywords are used to indicate different business scenarios, the business entry information is used to indicate the entry used to request text information inference, the application information is used to indicate the category of the application from which the target text information comes, and the intent description information is used to indicate the intent of the target text information; Obtaining, from the multiple system prompts, the system prompt corresponding to the determined business scenario.
3. The method according to claim 2, wherein The target text information is sourced from a picture; the determining the business scenario corresponding to the target text information according to one or more of the keyword, business entry information, application information, and intent description information related to the target text information includes: When the first picture corresponding to the target text information is requested for schedule extraction through a calendar application, if the first picture is obtained by taking a screenshot of the foreground application on the electronic device, determining the business scenario corresponding to the target text information according to the application information of the foreground application; If the first picture is not obtained by taking a screenshot on the electronic device, determining the intent description information corresponding to the first picture through a pre-trained intent classification model, and determining the business scenario corresponding to the target text information according to the intent description information corresponding to the first picture.
4. The method according to claim 1, wherein Before obtaining the KV value of the target system prompt from the cached KV values of the multiple system prompts, it further includes: For each system prompt in the multiple system prompts, inputting each system prompt into the target natural language model for processing to obtain the KV value of each system prompt; Caching the KV values of each system prompt in the multiple system prompts.
5. The method according to any one of claims 1 to 4, characterized in that, The obtaining target text information in response to a schedule extraction operation includes: In response to a schedule extraction operation on a first picture, a first recognition result and a second recognition result are determined through a target recognition model. The first recognition result includes the text recognition result of the first picture, and the second recognition result includes the edge recognition result and picture category of the first picture. The edge recognition result includes the graphic and text element attribute information or color block attribute information of the first picture. The target recognition model can determine the text recognition result, edge recognition result, and picture category of any picture; According to the picture category of the first picture, interference information unrelated to the schedule is filtered out from the text recognition result and the edge recognition result of the first picture; Based on the filtered first recognition result and the filtered second recognition result, the target text information is obtained.
6. The method according to claim 5, characterized in that The picture category includes instant messaging chat screenshots, instant messaging notification card screenshots, order screenshots, and other category pictures. The instant messaging chat screenshot refers to a picture obtained by taking a screenshot of a chat interface in an instant messaging application. The instant messaging notification card screenshot refers to a picture obtained by taking a screenshot of a service notification card in an instant messaging application. The order screenshot refers to a picture obtained by taking a screenshot of an order interface in an application program. The other category pictures include pictures other than the instant messaging chat screenshots, the instant messaging notification card screenshots, and the order screenshots; Among them, when the first picture is the instant messaging chat screenshot, the edge recognition result includes the graphic and text element attribute information. When the first picture is one of the instant messaging notification card screenshot, the order screenshot, and the other category pictures, the edge recognition result includes the color block attribute information.
7. The method according to claim 6, wherein The text recognition result includes text line coordinates; The filtering out of the interference information unrelated to the schedule from the text recognition result and the edge recognition result of the first picture according to the picture category of the first picture includes: When it is determined that the first picture is the instant messaging chat screenshot according to the picture category of the first picture, the recognition data corresponding to the skewed text lines in the text recognition result of the first picture is filtered out according to the text line coordinates of each text line in the text recognition result of the first picture; According to the text line coordinates of each text line in the text recognition result of the first picture, the recognition data corresponding to the text lines with a line height less than the target line height is filtered out from the text recognition result and the edge recognition result of the first picture. The target line height is the average line height of all text lines in the first picture; The recognition data corresponding to the chat timestamp in the first picture is filtered out from the text recognition result and the edge recognition result of the first picture.
8. The method according to claim 7, wherein The graphic and text element attribute information in the edge recognition result includes the coordinates and categories of chat elements, and the text recognition result also includes the text line recognition content; The filtering out of the recognition data corresponding to the chat timestamp in the first picture from the text recognition result and the edge recognition result of the first picture includes: Determine chat elements whose category is chat timestamp from the edge recognition results, and obtain at least one first candidate chat element; According to the coordinates of each first candidate chat element in the at least one first candidate chat element and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each first candidate chat element from the text recognition result of the first picture; When it is determined that the target first candidate chat element is a chat timestamp based on the text line recognition content corresponding to the target first candidate chat element, determine whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element, where the target first candidate chat element is any one of the at least one first candidate chat element; When it is determined to filter out the recognition data corresponding to the target first candidate chat element, filter out the recognition data corresponding to the target first candidate chat element from the text recognition result of the first picture and the edge recognition result.
9. The method according to claim 8, wherein The determining whether to filter out the recognition data corresponding to the target first candidate chat element according to the coordinates of the target first candidate chat element includes: When it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located in the middle position of the first picture and the area corresponding to the target first candidate chat element includes a single text line, determine to filter out the recognition data corresponding to the target first candidate chat element; or, When it is determined according to the coordinates of the target first candidate chat element that the target first candidate chat element is located on the right side of the first picture and the area corresponding to the target first candidate chat element includes a single text line, if the text line recognition content corresponding to the target first candidate chat element only includes a time point and does not include a date, determine to filter out the recognition data corresponding to the target first candidate chat element.
10. The method according to claim 6, characterized in that, The text recognition result includes text line coordinates and text line recognition content; The filtering out of interference information unrelated to the schedule from the text recognition result of the first picture and the edge recognition result according to the picture category of the first picture includes: When it is determined according to the category of the first picture that the first picture is a screenshot of an instant messaging notification card or an order screenshot, match the text line recognition content in each color block from the text recognition result of the first picture according to the coordinates of each color block in the edge recognition result and the text line coordinates in the text recognition result of the first picture; For any one of the color blocks, if it is determined according to the text line recognition content in the any one of the color blocks that the any one of the color blocks does not include content related to the schedule, filter out the recognition data corresponding to the any one of the color blocks from the edge recognition result and the text recognition result of the first picture.
11. The method according to any one of claims 6-10, characterized in that, The text recognition result includes text line coordinates and text line recognition content; The obtaining of the target text information based on the filtered first recognition result and the filtered second recognition result includes: Based on the coordinates of each object in the filtered edge recognition result and the text line coordinates in the text recognition result of the first picture, match the text line recognition content corresponding to each object from the text recognition result of the filtered first picture, where the object is a chat element or a color block; Determine the service scenario corresponding to the first picture according to the text line recognition content corresponding to each object and the picture category of the first picture; Based on the service scenario corresponding to the first picture, perform text splicing on the text line recognition content corresponding to each object to obtain the target text information.
12. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of claims 1-11 is implemented.
13. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is made to execute the method described in any one of claims 1-11.
Citation Information
Patent Citations
Information interaction method and device and readable storage medium
CN112527979A
Information recommendation method and device, electronic equipment and storage medium
CN113590743A
Cited By
Automatic schedule acquisition method for mobile operating system
CN121999475A
Automatic schedule acquisition method for mobile operating system
CN121999475B