Data processing method and device, electronic equipment and computer readable medium

By combining user input information and screen display content with tag recognition technology, target text is generated and tasks are determined, solving the problem of insufficient understanding of intent in fuzzy query scenarios by traditional interaction methods, realizing intelligent one-click screen inquiry and improving the accuracy of interaction.

CN122640501APending Publication Date: 2026-08-25GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510215747.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Traditional interaction methods struggle to understand a user's true intent in fuzzy query scenarios, especially in the absence of clear context, resulting in a poor interactive experience, insufficient integration and utilization of cross-modal information, and an inability to achieve truly intelligent one-click screen inquiry.

Method used

By acquiring user input information and tags of the content currently displayed on the electronic device screen, and combining technologies such as optical character recognition, layout understanding, subject recognition, scene category recognition, entity recognition, and image and text description, the system identifies the content currently displayed on the screen, generates target text, and determines the target task based on a large language model.

Benefits of technology

It improves the accuracy of user interaction, enabling it to more accurately understand users' ambiguous intentions, execute target tasks that are closer to users' actual needs, and achieve intelligent one-click screen inquiry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640501A_ABST
    Figure CN122640501A_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and device, electronic equipment and computer readable medium, and belongs to the technical field of data processing. The method comprises the following steps: obtaining input information of a user; obtaining a label corresponding to current display content of a screen of the electronic equipment; performing screen recognition on the current display content of the screen to obtain target text for describing the current display content of the screen in the case that the input information is associated with the label; determining a target task based on the input information and the target text; and executing the target task. On the basis of screen recognition, the target task is determined in combination with the input information of the user and the display content of the current screen. For the fuzzy query or fuzzy intention of the user, the intention of the user can be more accurately understood, and the target task closer to the actual demand of the user can be executed, so that the accuracy of interaction is improved, and intelligent and non-sensing interaction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a data processing method, apparatus, electronic device, and computer-readable medium. Background Technology

[0002] In recent years, with the widespread adoption of smartphones and voice assistants, users' interaction needs on mobile devices have become increasingly diverse and complex. Among these, the frequency with which users ask fuzzy queries (such as "What is this?") in specific scenarios is gradually increasing. However, these fuzzy queries often lack clear context, making it difficult for traditional interaction methods to understand the user's true intent. Summary of the Invention

[0003] This application proposes a data processing method, apparatus, electronic device, and computer-readable medium to improve the above-mentioned deficiencies.

[0004] In a first aspect, this application provides a data processing method applied to an electronic device, the method comprising: acquiring user input information; acquiring a tag corresponding to the currently displayed content on the screen of the electronic device; performing screen recognition on the currently displayed content on the screen, when the input information is associated with the tag, to obtain target text describing the currently displayed content on the screen; determining a target task based on the input information and the target text; and executing the target task.

[0005] Secondly, this application also provides a data processing apparatus applied to an electronic device, the apparatus comprising: a first acquisition unit for acquiring user input information; a second acquisition unit for acquiring a tag corresponding to the currently displayed content on the screen of the electronic device; a recognition unit for performing screen recognition on the currently displayed content on the screen, when the input information is associated with the tag, to obtain target text describing the currently displayed content on the screen; a determination unit for determining a target task based on the input information and the target text; and an execution unit for executing the target task.

[0006] Thirdly, this application also provides an electronic device, comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory, the one or more application programs are configured to be executed by the one or more processors, and the one or more application programs are configured to perform the methods described above.

[0007] Fourthly, this application also provides a computer-readable medium storing processor-executable program code that, when executed by the processor, causes the processor to perform the above-described method.

[0008] The solution provided in this application first obtains user input information; then obtains a tag corresponding to the currently displayed content on the screen of the electronic device; then, when the input information is associated with the tag, screen recognition is performed on the currently displayed content to obtain target text describing the currently displayed content; secondly, a target task is determined based on the input information and the target text; and finally, the target task is executed.

[0009] In this embodiment, based on the need for screen recognition, the target task is determined by combining the user's input information and the current screen display content. For the user's fuzzy query or fuzzy intent, the user's intent can be understood more accurately, and the target task that is closer to the user's actual needs can be executed, thereby improving the accuracy of the interaction.

[0010] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart of a data processing method provided in an embodiment of this application is shown;

[0013] Figure 2 A flowchart of a data processing method according to another embodiment of this application is shown;

[0014] Figure 3 A flowchart of a data processing method provided in another embodiment of this application is shown;

[0015] Figure 4 A flowchart of a data processing method according to another embodiment of this application is shown;

[0016] Figure 5 A structural block diagram of the data processing apparatus provided in an embodiment of this application is shown;

[0017] Figure 6 A structural block diagram of a data processing apparatus provided in another embodiment of this application is shown;

[0018] Figure 7A structural block diagram of the electronic device provided in an embodiment of this application is shown;

[0019] Figure 8 A structural block diagram of a computer-readable storage medium provided in an embodiment of this application is shown. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.

[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0022] In recent years, with the widespread adoption of smartphones and voice assistants, users' interaction needs on mobile devices have become increasingly diverse and complex. Among these, the frequency with which users ask fuzzy queries (such as "What is this?") in specific scenarios is gradually increasing. However, these fuzzy queries often lack clear context, making it difficult for traditional interaction methods to understand the user's true intent.

[0023] Although the rapid development of large models in recent years has given mobile phone assistants stronger natural language processing capabilities, the following shortcomings exist in specific application scenarios:

[0024] 1. Insufficient accuracy in intent judgment: In fuzzy query scenarios, due to the lack of sufficient contextual information, voice assistants have difficulty accurately identifying the user's true needs, requiring the user to continuously supplement the description of the needs, resulting in a poor user interaction experience.

[0025] 2. Lack of cross-modal capabilities: Most current mobile smart assistants mainly rely on text information for intent understanding, and there is insufficient integration and utilization of cross-modal information based on the user's current screen environment (such as the app being used, screen image tags, etc.), making it difficult to achieve dynamic response and accurate judgment.

[0026] 3. Limitations of screen recognition interaction: Although some smart devices support functions such as screen asking, they are mostly based on user-initiated behaviors, such as manual selection and dragging, to trigger screen recognition interaction. This fails to effectively combine screen content with user intent and cannot achieve true intelligent one-click screen asking.

[0027] While some manufacturers' voice assistants currently support cross-modal information interaction, these are mostly based on user-initiated interactions. This means the user has explicitly indicated they want to interact based on selected screen content, photos, or documents, among other modal data. For example, if a user wants to access an image from an app they're browsing, they need to drag the image into the interaction box and then enter the query: "Where is this?". Or, if a user wants to call a contact in their photo album, they need to circle the phone number on the screen and then enter the query: "Call this number." While these solutions achieve multimodal information understanding, they don't truly achieve one-click cross-modal user intent recognition.

[0028] As users' demands for intelligent assistant functions continue to rise, these problems are becoming increasingly prominent, especially in scenarios that support real-time screen navigation, shopping recommendations, and travel planning. The lack of cross-modal fuzzy intent understanding leads to a poor user experience. For example, in a voice assistant, if a user queries "navigate here," the system cannot provide a reasonable response if it cannot analyze the user's current app status and screen content.

[0029] Therefore, in order to overcome the above-mentioned defects, embodiments of this application provide a data processing method applied to an electronic device, such as a smartphone or tablet computer. Figure 1 As shown, the method may include S101 to S105.

[0030] S101: Obtain user input information.

[0031] It is understood that the input information refers to the user's request or instruction information input to the electronic device, which can be input into the electronic device in different forms. For example, the request or instruction information can be transmitted to the electronic device in the form of voice, text, images, videos, documents, and photos.

[0032] It should be noted that after obtaining the user's input information, the input information is converted into structured data to facilitate understanding and processing by downstream modules.

[0033] In one implementation, after a user opens the target application on an electronic device, they can ask "What is this?" by voice while browsing a webpage. The electronic device can then convert the user's voice into text information through a voice recognition module, which serves as the user's input. It should be noted that the target application is an application that executes the method of the embodiments of this application. Opening the target application allows the method of this application to be executed.

[0034] As another implementation method, after a user opens the target application on an electronic device and browses pictures of tourist attractions, they can issue a voice command such as "navigate to here". The electronic device can then convert the voice recognition module into text information, which can be used as the user's input information.

[0035] As another implementation method, after the user opens the target application on the electronic device, they can input their needs or instructions via text as user input.

[0036] S102: Obtain the tag corresponding to the currently displayed content on the screen of the electronic device.

[0037] It is understood that the displayed content refers to the content contained within the interface of the screen. The screen interface can be the interface that the user is currently browsing through the screen. The displayed content can be images, videos, text, etc. within the interface. Specifically, the text can be text entered by the user or text converted from voice input and displayed on the interface, or text on the webpage that the user is browsing. The images and videos can be content that the user is browsing.

[0038] It should be noted that this label refers to the type of displayed content. Specifically, it can refer to the overall type of displayed content, such as the label for the currently displayed content being wallpaper, chat interface, or WeChat screenshot. It can also refer to the type of a specific object within the displayed content, such as the label for the currently displayed content being beach, waves, or sunset.

[0039] One implementation method is to obtain the currently displayed content of the electronic device screen after the target application is running on the user's screen. For example, after the user opens the target application, a screenshot of the screen interface can be taken, and the screenshot image can be identified to obtain the displayed content. Alternatively, a screenshot of the screen interface can be taken after detecting that the user has performed a specified operation within the screen interface. This specified operation may include the user triggering a specified control, which is used for the user to input query content within the screen interface. For example, the user can input text or voice, which is considered a triggered specified control.

[0040] As another implementation, the current content displayed on the electronic device's screen can be retrieved after the target application is run on the user's screen and the user's input information is obtained. For example, after the user opens the target application and the user's input information is obtained, a screenshot of the screen interface can be retrieved, and the screenshot can be recognized to obtain the displayed content.

[0041] Obtaining the tag corresponding to the currently displayed content on the screen of an electronic device can be achieved by: determining the application corresponding to the currently displayed content; if the application is a preset application, then using the preset tag corresponding to the application as the tag corresponding to the currently displayed content. If the application is not a preset application, then obtaining the currently displayed content on the screen of the electronic device; identifying the target object of the currently displayed content; and determining the tag corresponding to the currently displayed content based on the target object. For details, please refer to subsequent embodiments.

[0042] S103: When the input information is associated with the label, screen recognition is performed on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen.

[0043] Understandably, users input various information into electronic devices, such as opening music apps, connecting via Bluetooth, navigating here, asking "What is this?", or booking hotels. Furthermore, users are increasingly making fuzzy queries (e.g., "What is this?") in specific scenarios. For input information with ambiguous intent, it's necessary to access the currently viewed screen and extract key information from it. Only based on this key information and the input information can the user's intent be clarified. Therefore, it's necessary to determine whether key information needs to be extracted from the currently displayed screen content.

[0044] It should be noted that if the user's input information is associated with the label, it means that the user's needs or instructions are related to the current content displayed on the screen. In this case, it is necessary to determine the user's true intention based on the current content displayed on the screen.

[0045] As one embodiment, it is determined whether the input information contains the tag. If the tag is present, the input information is associated with the tag. For example, if the user's input information is "phone model," and the tag corresponding to the currently displayed content is "phone," it indicates that the currently displayed content is related to a phone. Therefore, it is determined that the input information is associated with the tag, and screen recognition of the currently displayed content is required. The user's true intention is likely to ask about the model of the currently displayed phone, not the model of the electronic device (phone). If the tag corresponding to the currently displayed content is "beach," it indicates that the currently displayed content is related to a beach. Therefore, it is determined that the input information is not associated with the tag, and the user's true intention is likely to ask about the model of the electronic device (phone). In this case, screen recognition of the currently displayed content is not required. Using tags can more accurately determine the user's true intention and has a disambiguation effect on ambiguous scenarios in one-click screen queries.

[0046] As another embodiment, the association between the input information and the tag can be determined by the specific content of the input information. For example, if the input information includes an indicator pronoun or an interrogative pronoun, it is determined that the input information is associated with the tag; if the input information does not include an indicator pronoun or an interrogative pronoun, it is determined that the input information is not associated with the tag. For details, please refer to the following embodiments.

[0047] It should be noted that screen recognition can be performed on the currently displayed content of the screen by using at least one of Optical Character Recognition (OCR), layout understanding, subject recognition, scene category recognition, entity recognition, and image and text description to obtain target text used to describe the currently displayed content of the screen.

[0048] As one implementation method, screen recognition can be performed using optical character recognition (OCR). For example, after obtaining a screenshot of the current screen, OCR can be performed on the screenshot to obtain the text information in the screenshot.

[0049] As another implementation method, screen recognition can be performed through panel understanding. For example, after obtaining a screenshot of the current screen, version analysis is performed on the screenshot to obtain information such as the position, layout, and font size of each text segment.

[0050] As another implementation method, screen recognition can be performed through subject recognition. For example, after obtaining a screenshot of the current screen, the subject of the screenshot is determined. For example, if the determination result is a face, a scenic spot, or a vehicle, the relevant subject result is output as a person, scenic spot name, or vehicle model.

[0051] As another implementation method, screen recognition can be performed through scene category recognition. For example, after obtaining a screenshot of the current screen, scene category recognition is performed on the screenshot. The current screenshot is classified through image classification and tag algorithms, and multi-tag content such as documents, tables, test papers and images is output.

[0052] As another implementation method, screen recognition can be performed through entity recognition. For example, after obtaining a screenshot of the current screen, entity recognition can be performed on the screenshot. Based on the results of optical character recognition, an entity recognition algorithm can be introduced to output entity results such as address, name, time, and telephone number.

[0053] As another implementation method, screen recognition can be performed through graphic and textual descriptions. For example, after obtaining a screenshot of the current screen, a graphic and textual description of the screenshot is performed. By introducing a multimodal model, the current screen image is described and summarized in multiple modes.

[0054] It should be noted that screen recognition can be performed on the currently displayed content using at least one of the following methods: optical character recognition, layout understanding, subject recognition, scene category recognition, entity recognition, and image-text description, in any combination, to obtain target text describing the currently displayed content. The target text is then converted into a format that can be directly used by the large language model, facilitating its subsequent operation.

[0055] S104: Determine the target task based on the input information and the target text.

[0056] It is understandable that this objective task represents the task that the large language model needs to perform based on the input information and the target text, and this objective task corresponds to the user's input information. For example, this objective task could be to open navigation software and navigate to a certain location.

[0057] The large language model can determine the target task based on the input information and the target text. That is, the input information and the text information used to describe the content currently displayed on the screen are input into the large language model, and the large language model can determine the target task.

[0058] It should be noted that the large language module includes a Function Call screen recognition intent judgment module, a screen recognition intent execution module, and a Function Call API understanding module. The Function Call screen recognition intent judgment module is mainly used to determine whether screen recognition is needed. The screen recognition intent execution module is used for specific decision-making actions, such as performing screen recognition or not performing screen recognition. The Function Call API understanding module is mainly used to obtain the target task, as well as the corresponding API interface and API slot information based on the input information and target text. For details, please refer to the subsequent embodiments.

[0059] Furthermore, if the input information is not associated with the label, the target task is determined based on the input information.

[0060] If the input information is not associated with the label, it means the user's request is unrelated to the currently displayed content on the screen, and executing the user's request does not require accessing the currently displayed content. Therefore, the target task can be determined directly based on the input information using the Function Call API understanding module.

[0061] For example, if the user's input information is "call Mr. Wang", and the label of the currently displayed content on the screen is "wallpaper", and it is determined that the input information "call Mr. Wang" is not associated with the label "wallpaper", then based on the input information, the target task is determined to be making a phone call, and the call recipient is Mr. Wang.

[0062] S105: Execute the target task.

[0063] Understandably, the target task depends on the user's input information. If the user's input information is "connect to Bluetooth", then the target task is to connect to Bluetooth. If the user's input information is "navigate to here", then the target task is to navigate. Executing the target task is to fulfill the user's actual needs.

[0064] In one implementation, when a user opens the target application, the user says "Where is this?" and the user's input information is obtained as "Where is this?". The label of the content currently displayed on the screen of the electronic device is obtained as "attraction". With the input information and the label associated, screen recognition is performed on the content currently displayed on the screen to obtain target text describing the content currently displayed on the screen. The target task is determined based on the input information and the target text, and then the target task is executed.

[0065] In this embodiment, based on the need for screen recognition, the target task is determined by combining the user's input information and the current screen display content. For the user's fuzzy query or fuzzy intent, the user's intent can be understood more accurately, and the target task that is closer to the user's actual needs can be executed, thereby improving the accuracy of the interaction.

[0066] Please see Figure 2 This application provides a data processing method, which is applied to the above-mentioned electronic device. Specifically, the method includes: S201 to S209.

[0067] S201: Obtain user input information.

[0068] S202: Determine the application corresponding to the currently displayed content on the screen of the electronic device.

[0069] It is understandable that the application is the one currently running on the screen of the electronic device. For example, the user is browsing a travel guide article by a blogger on the Xiaohongshu APP, meaning that the application corresponding to the content currently displayed on the screen is Xiaohongshu.

[0070] For example, the application can be the APP name or the APP package name, with the APP package name being the application's identification code.

[0071] It should be noted that when using an electronic device, a user may open the Xiaohongshu app and then the Weather app. At this time, the electronic device is running two applications in the background: the Xiaohongshu app and the Weather app. If the current screen displays a Xiaohongshu webpage, then the application corresponding to the current screen is determined to be the Xiaohongshu app; if the current screen displays a Weather webpage, then the application corresponding to the current screen is determined to be the Weather app.

[0072] S203: If the application is a preset application, then the preset tag corresponding to the application is used as the tag corresponding to the currently displayed content on the screen.

[0073] It is understandable that the preset application is a pre-set application. For example, if the preset application is the desktop system, it means that the application corresponding to the current display content on the screen is the desktop system. If the preset label corresponding to the desktop system is "wallpaper", then "wallpaper" will be used as the label corresponding to the current display content on the screen.

[0074] S204: If the application is not a preset application, then obtain the current display content of the screen of the electronic device.

[0075] Understandably, if the application is not a pre-installed application, it needs to determine the label based on the specific content displayed on the screen, which requires obtaining the current content displayed on the screen. Specifically, triggering a screenshot operation on that screen and using the screenshot image as the current content displayed on the screen.

[0076] S205: Identify the target object of the content currently displayed on the screen.

[0077] It is understandable that target recognition of the currently displayed content on the screen can be performed to obtain the target object of the displayed content. This can be done using image processing methods, machine learning methods, or deep learning methods to identify the target image of the currently displayed content on the screen.

[0078] It should be noted that when the current content displayed on the screen contains multiple target objects, the primary target object is identified, and the primary target object includes at least one.

[0079] In one embodiment, the current content displayed on the screen is an electric car driving on a road by the sea. The current content displayed on the screen is obtained by taking a screenshot. The target objects are obtained by object detection on the image, which are the car and the road.

[0080] In another embodiment, the content currently displayed on the screen is a WeChat chat dialog box. A screenshot image is obtained by taking a screenshot of the currently displayed content on the screen. Object detection is performed on the image to obtain the target objects as the dialog box and text.

[0081] S206: Determine the label corresponding to the currently displayed content on the screen based on the target object.

[0082] Understandably, different target objects correspond to different labels. For example, a car driving on a road by the sea: if the car body occupies a large proportion of the entire screenshot image, its corresponding label is "car"; if the car body occupies a small proportion of the entire screenshot image, and the road occupies a large proportion, its corresponding label is "road." If ocean waves occupy a large proportion of the entire screenshot image, its corresponding label is "waves."

[0083] Therefore, it is necessary to determine the label of the currently displayed content on the screen based on the target object.

[0084] S207: When the input information is associated with the label, screen recognition is performed on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen.

[0085] S208: Determine the target task based on the input information and the target text.

[0086] S209: Execute the target task.

[0087] Please see Figure 3 This application provides a data processing method, which is applied to the above-mentioned electronic device. Specifically, the method includes: S301 to S307.

[0088] S301: Obtain user input information.

[0089] S302: Obtain the tag corresponding to the currently displayed content on the screen of the electronic device.

[0090] S303: If the input information includes an indicator pronoun or an interrogative pronoun, then the input information is determined to be associated with the tag.

[0091] It is understandable that the user's vague input information needs to be combined with the currently displayed content on the screen to obtain the target task, while some users' accurate input information does not need to be combined with the currently displayed content on the screen, and the target task can be obtained directly based on the user's input information. Therefore, it is necessary to analyze and judge the user's input information to determine whether the input information is associated with the label.

[0092] It should be noted that if the input information includes demonstrative pronouns or interrogative pronouns, it indicates that the input information is a vague intent or vague instruction, and thus the input information is associated with punctuation. Demonstrative pronouns include those referring to people or things, those indicating specificity or choice, and those indicating location, such as "this," "these," "here," "that," "those," and "there." Interrogative pronouns include: "what," "which," "how many," "where," "why," and "who," etc.

[0093] In one implementation, if the input information is "What is this", "Where is this", "Who is on the mountain", or "Navigate here", it means that the input information includes an indicator pronoun or interrogative pronoun. In this case, it is determined that the input information is associated with the label, and screen recognition of the currently displayed content is required.

[0094] S304: If the input information does not include indicator pronouns and interrogative pronouns, then it is determined that the input information is not associated with the label.

[0095] Understandably, if the input information does not include demonstrative pronouns or interrogative pronouns, it means that the user's input information is clear, and therefore it is determined that the input information is not associated with the label.

[0096] In one implementation, if the input information is "connect Bluetooth" or "play the music 'Tomorrow Will Be Better'", then the input information does not include demonstrative pronouns or interrogative pronouns, which means that the input information is not associated with the tag.

[0097] Furthermore, it is determined whether the input information includes the tag. If the input information includes the tag, it is determined that the input information is associated with the tag; if the input information does not include the tag, it is determined whether the input information includes an indicator pronoun or an interrogative pronoun. If the input information includes an indicator pronoun or an interrogative pronoun, it is determined that the input information is associated with the tag; if the input information does not include an indicator pronoun or an interrogative pronoun, it is determined that the input information is not associated with the tag. This allows for a more accurate understanding of user intent, thereby enabling more accurate responses to user needs.

[0098] S305: When the input information is associated with the label, screen recognition is performed on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen.

[0099] S306: Determine the target task based on the input information and the target text.

[0100] S307: Execute the target task.

[0101] Please see Figure 4 This application provides a data processing method, which is applied to the above-mentioned electronic device. Specifically, the method includes: S401 to S406.

[0102] S401: Obtain user input information.

[0103] S402: Obtain the tag corresponding to the currently displayed content on the screen of the electronic device.

[0104] S403: Determine whether the electronic device has screen recognition function based on the model and system version of the electronic device.

[0105] Understandably, as mobile phone smart assistants continue to develop, different mobile phone manufacturers are developing their own mobile phone smart assistants. Moreover, as mobile phone versions are constantly updated, older versions of mobile phones may lack some hardware components, so some versions of mobile phones cannot realize screen recognition functions.

[0106] Based on the model and system version of the electronic device, determine whether the electronic device has screen recognition functionality. Specifically, if the model of the electronic device has screen recognition hardware and the operating system corresponding to the system version has the function of calling screen recognition, then the electronic device is determined to have screen recognition functionality.

[0107] S404: If screen recognition function is available, screen recognition is performed on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen.

[0108] If the device has screen recognition functionality, it means that both the hardware and software of the electronic device meet the requirements for screen recognition. Therefore, screen recognition can be performed on the currently displayed content of the screen to obtain target text that describes the currently displayed content of the screen.

[0109] S405: Determine the target task based on the input information and the target text.

[0110] It is understandable that the target task determined based on a large language model may contain inaccurate information. Therefore, the authenticity of the target task can be verified. Specifically, this can be done by determining the corresponding API service interfaces and API slots based on the target task, and then validating these API service interfaces and API slots to achieve the effect of verifying the authenticity of the target task.

[0111] Because of the complexity of cross-modal information input in intelligent dialogue systems, the results often lead to the model's "fantasy" input. For example, the address, phone number, or name of the slot may not actually exist in the input. Therefore, API result validation is necessary to prevent the use of fictitious slot information, which could result in negative user reviews and complaints. Examples include dialing fictitious phone numbers or navigating to non-existent addresses.

[0112] Specifically, API verification mainly includes three steps: API name verification, API slot name verification, and API slot value verification.

[0113] It's important to know that the API documentation includes the name of each API interface and API slot in the electronic device. You can refer to the API documentation to determine if the API name and API slot name corresponding to the target task are correct. By checking if the API slot value is included in the target text information extracted across modalities, if it is, the API slot value is correct; otherwise, it is incorrect. In this case, it's necessary to return to the point where the input information is associated with the tag, perform screen recognition on the currently displayed content, obtain the target text describing the currently displayed content, and then proceed with subsequent operations.

[0114] S406: Execute the target task.

[0115] Please see Figure 5 The diagram shows a structural block diagram of a data processing device 500 provided in an embodiment of this application. The testing device 500 includes: a user input module 511, a mobile terminal module 512, a mobile phone screen module 513, a cross-modal understanding module 514, an intelligent dialogue system module 515, and an output legality verification module 516.

[0116] User input module 511 is used to acquire user input information and transmit the received input information to intelligent dialogue system module 515.

[0117] The mobile terminal module 512 is used to obtain the APP package name (APP identity information) corresponding to the currently displayed content on the screen, and transmit the current APP package name to the intelligent dialogue system module 515. The mobile terminal module 512 is also used to obtain other terminal information of the electronic device (electronic device model and system version information), and transmit the other terminal information of the electronic device to the cross-modal understanding module 514.

[0118] The mobile phone screen module 513 is used to obtain the terminal screen tag (the tag of the screenshot image of the currently displayed content on the screen) and the current screen image (the screenshot image of the currently displayed content on the screen), and transmits the obtained terminal screen tag to the intelligent dialogue system module 515, and transmits the current screen image to the cross-modal understanding module 514.

[0119] The intelligent dialogue system module 515 determines the screen recognition intent based on the acquired user input information, the current APP package name, and the terminal screen tags. If the result indicates that screen recognition is required, the cross-modal understanding module 514 performs cross-modal understanding of the currently displayed content to obtain the target text. Then, based on a large language model, the target text and input information are processed to obtain the target task. If the result indicates that screen recognition is not required, the large language model and user input information determine the target task, resulting in the corresponding API selection and slot understanding. The output validity verification module 516 verifies the API; if the verification is successful, the target task is executed.

[0120] The cross-modal understanding module 514 is used to perform cross-modal understanding of the content currently displayed on the screen, including OCR recognition, layout understanding, subject recognition, scene classification, entity recognition, and image and text description.

[0121] Module 515 of the intelligent dialogue system, also known as the large language module, includes a Function Call screen recognition intent judgment module, a screen recognition intent execution module, and a Function Call API understanding module.

[0122] Based on the Function Call capability of the Large Language Model (LLM), the system performs final processing on user input and cross-modal input and feeds back the final execution API to the user. The intelligent dialogue system module 515 is divided into three parts: Function Call screen recognition intent judgment module, screen recognition intent execution module, and Function Call API understanding module.

[0123] The Function Call screen recognition intent determination module receives user input and endpoint information (the application corresponding to the currently displayed content on the screen), then makes a decision and outputs an indicator indicating whether screen recognition is needed. For example, based on the input query "What is this?", the current app package name is "Xiaohongshu", and the endpoint tag (the label corresponding to the currently displayed content on the screen) is "phone". In this case, the intent of the query information is very vague and lacks sufficient context. However, the endpoint information indicates that the user's intent requires the use of screen information. In this case, the model will decide to output the intent to recognize the screen. If the input query is "Turn on this Bluetooth", the current app package name is "System Desktop", and the endpoint tag is "Wallpaper", the user's intent is very clear and does not need to refer to screen information to satisfy the user's intent. In this case, the model decides that screen recognition is not necessary.

[0124] The screen recognition intent execution module handles screen recognition identifiers. If screen recognition is required, it invokes the cross-modal understanding module to understand and analyze the current screen image. If screen recognition is not required, it directly passes the user input and endpoint information from the previous module to the next module.

[0125] The Function Call API Understanding module: If screen recognition is triggered, it receives user input, endpoint information, and cross-modal input, makes a decision, selects the appropriate API, and provides the corresponding API slot information. For example, if the user query is "Navigate here", the current app package name is "WeChat", and the endpoint tag is "WeChat screenshot", and the previous module has already triggered screen recognition and there is cross-modal information input, this module needs to select the most suitable API: "route_navigation" based on all the input information, then understand the specific content on the screen, and output the correct navigation location, `route_navigation(to_location = Starbucks on Yunnan Road, Shanghai)`, thus achieving a complete one-click screen query intent recognition. When the previous module has not triggered screen recognition and there is no cross-modal information input, this module processes plain text modal input requests normally. For example, if the user query is "Navigate to Taikoo Li", the current app package name is "WeChat", and the endpoint tag is "WeChat screenshot", the final output will be `route_navigation(to_location = Taikoo Li)`.

[0126] This application leverages LLM-based complex semantic understanding capabilities to analyze user queries, mobile device status information, and device-side screen tag information in real time. It automatically identifies whether the user's current intent requires recognition and understanding of screen information. For example, when a user says "Navigate here," screen recognition is automatically triggered based on the current user context, device status information (APP, etc.), and device-side screen image tags. The image content is then analyzed, and the final interactive response for the location on the screen is output, thus achieving truly intelligent and seamless interaction.

[0127] Please see Figure 6 The diagram shows a structural block diagram of a data processing device 600 provided in an embodiment of this application. The testing device 600 includes: a first acquisition unit 610, a second acquisition unit 620, an identification unit 630, a determination unit 640, and an execution unit 650.

[0128] The first acquisition unit 610 is used to acquire user input information.

[0129] The second acquisition unit 620 is used to acquire the tag corresponding to the currently displayed content on the screen of the electronic device.

[0130] Furthermore, the second acquisition unit 620 can also be used to determine the application corresponding to the currently displayed content on the screen of the electronic device; if the application is a preset application, then the preset tag corresponding to the application is used as the tag corresponding to the currently displayed content on the screen.

[0131] Furthermore, the second acquisition unit 620 can also be used to acquire the currently displayed content of the screen of the electronic device if the application is not a preset application; identify the target object of the currently displayed content of the screen; and determine the tag corresponding to the currently displayed content of the screen based on the target object.

[0132] Furthermore, the second acquisition unit 620 can also be used to determine that the input information is associated with the tag if the input information includes an indicator pronoun or an interrogative pronoun; and to determine that the input information is not associated with the tag if the input information does not include an indicator pronoun or an interrogative pronoun.

[0133] The recognition unit 630 is used to perform screen recognition on the currently displayed content of the screen when the input information is associated with the label, so as to obtain target text describing the currently displayed content of the screen.

[0134] Furthermore, the identification unit 630 can also be used to determine whether the electronic device has a screen recognition function based on the model and system version of the electronic device; if it has a screen recognition function, it performs screen recognition on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen.

[0135] Furthermore, the recognition unit 630 can also be used to perform at least one of optical character recognition, layout understanding, subject recognition, scene category recognition, entity recognition, and graphic description on the currently displayed content of the screen to obtain target text for describing the currently displayed content of the screen.

[0136] The determining unit 640 is used to determine the target task based on the input information and the target text.

[0137] Furthermore, the determining unit 640 can also be used to determine the target task based on the input information when the input information is not associated with the label.

[0138] The execution unit 650 is used to execute the target task.

[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0140] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.

[0141] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0142] Please refer to Figure 7 This diagram illustrates a structural block diagram of an electronic device 700 provided in an embodiment of this application. The electronic device 700 can be an in-vehicle infotainment system, which can be installed in a vehicle. The electronic device 700 in this application may include one or more of the following components: a processor 711, a memory 712, and one or more application programs, wherein the processor 711 is electrically connected to the memory 712, and the one or more programs are configured to execute the methods described in the foregoing embodiments of the test methods.

[0143] The processor 711 may include one or more processing cores. The processor 711 connects to various parts within the electronic device 700 using various interfaces and lines, and performs various functions and processes data of the electronic device 700 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 712, and by calling data stored in the memory 712. Optionally, the processor 711 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 711 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and computer programs; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 711 and may be implemented separately using a communication chip. Specifically, the methods described in the foregoing embodiments can be executed by one or more processors 711.

[0144] In some implementations, memory 712 may include random access memory (RAM) or read-only memory (ROM). Memory 712 can be used to store instructions, programs, code, code sets, or instruction sets. Memory 712 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the electronic device 700 during use.

[0145] Please refer to Figure 8 This diagram illustrates a structural block diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 800 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0146] The computer-readable medium 800 may be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable medium 800 includes a non-transitory computer-readable storage medium. The computer-readable medium 800 has storage space for program code 810 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 810 may be compressed, for example, in a suitable form.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A data processing method, characterized in that, Applied to electronic devices, the method includes: Obtain user input information; Obtain the tag corresponding to the currently displayed content on the screen of the electronic device; When the input information is associated with the label, screen recognition is performed on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen; The target task is determined based on the input information and the target text; Perform the target task.

2. The method according to claim 1, characterized in that, Also includes: If the input information is not associated with the label, the target task is determined based on the input information.

3. The method according to claim 1, characterized in that, The step of obtaining the tag corresponding to the currently displayed content on the screen of the electronic device includes: Determine the application corresponding to the currently displayed content on the screen of the electronic device; If the application is a preset application, then the preset tag corresponding to the application will be used as the tag corresponding to the currently displayed content on the screen.

4. The method according to claim 3, characterized in that, Also includes: If the application is not a preset application, then obtain the current screen display content of the electronic device; Identify the target object of the content currently displayed on the screen; The label corresponding to the currently displayed content on the screen is determined based on the target object.

5. The method according to claim 1, characterized in that, After obtaining the tag corresponding to the currently displayed content on the screen of the electronic device, the method further includes: If the input information includes an indicator pronoun or an interrogative pronoun, then the input information is determined to be associated with the tag; If the input information does not include demonstrative pronouns and interrogative pronouns, then it is determined that the input information is not associated with the tag.

6. The method according to claim 1, characterized in that, The step of performing screen recognition on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen includes: Based on the model and system version of the electronic device, determine whether the electronic device has screen recognition functionality; If screen recognition is available, screen recognition is performed on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen.

7. The method according to claim 1, characterized in that, The step of performing screen recognition on the currently displayed content of the screen to obtain target text describing the currently displayed content of the screen includes: Perform at least one of optical character recognition, layout understanding, subject recognition, scene category recognition, entity recognition, and image and text description on the currently displayed content of the screen to obtain target text for describing the currently displayed content of the screen.

8. A data processing apparatus, characterized in that, Applied to electronic devices, the device includes: The first acquisition unit is used to acquire user input information; The second acquisition unit is used to acquire the tag corresponding to the currently displayed content on the screen of the electronic device; The recognition unit is configured to perform screen recognition on the currently displayed content of the screen when the input information is associated with the label, and obtain target text describing the currently displayed content of the screen. A determining unit is configured to determine a target task based on the input information and the target text; An execution unit is used to execute the target task.

9. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory, the one or more applications are configured to be executed by the one or more processors, and the one or more applications are configured to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-7.