Data processing method, electronic device, readable medium and program product
By providing local data and parameter adjustment controls in the generative AI application interface, the problem of users needing to input professional prompts is solved, resulting in generated content that better meets user expectations and improving the user experience.
Patent Information
- Application Number
- PCT/CN2025/096938
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-10
- Filing Date
- 2025-05-23
- Publication Date
- 2026-02-05
AI Technical Summary
Existing applications based on generative AI models require users to input specialized prompts, resulting in a significant gap between the generated content and user expectations, leading to a poor user experience.
By providing controls and parameter tuning controls that relate to local data in the application interface, users can set control parameters, and the generation model can generate personalized content that meets the user's needs based on these parameters.
It improves the user experience, makes the generated content more accurately meet user expectations, reduces users' reliance on professional prompts, and enhances the application's personalization and convenience.
Smart Images

Figure CN2025096938_05022026_PF_FP_ABST
Abstract
Description
Data processing methods, electronic devices, readable media and program products
[0001] This application claims priority to Chinese Patent Application No. 202411060025.0, filed on August 2, 2024, entitled “Data Processing Method, Electronic Device, Readable Medium and Program Product”, the entire contents of which are incorporated herein by reference.
[0002] Furthermore, this application claims priority to Chinese Patent Application No. 202411271842.0, filed on September 10, 2024, entitled "Data Processing Method, Electronic Device, Readable Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention relates to the field of computer technology, and more specifically to a data processing method, electronic device, readable medium, and program product. Background Technology
[0004] Generative artificial intelligence (AIGC) is an important branch of artificial intelligence (AI). Based on generative adversarial networks (GANs) and large pre-trained language models (hereinafter referred to as "large models"), AIGC can generate content with a certain degree of creativity and quality using artificial intelligence algorithms. For example, AIGC can generate matching text, images, audio, and video content based on input keywords, prompts, or descriptive samples.
[0005] However, currently, large models (also known as artificial intelligence models or AI models) developed based on AIGC technology provide functions such as text expansion and AI drawing through specific application interfaces (hereinafter referred to as application interfaces). These functions require users to input relatively specialized prompts as model input. This results in the generated content often deviating significantly from the user's expectations when using these specific applications developed based on large models, leading to a poor user experience. Summary of the Invention
[0006] This application provides a data processing method, electronic device, readable medium, and program product. Users can set control parameters for specified text, images, audio, video, and other target content on the relevant interface of the target application. These control parameters can be associated with relevant locally stored data or descriptive information of the relevant data, and can ultimately control the generation of personalized content that accurately meets the user's needs, thereby improving the user experience.
[0007] In a first aspect, this application provides a data processing method applied to an electronic device. The method includes: displaying a first interface of a first application, wherein the first interface includes first display content; detecting a first processing instruction for the first display content; processing the first display content according to reference data to generate a first processing result, wherein the first processing result includes information obtained according to the reference data; and displaying the first processing result on the first interface.
[0008] For example, the aforementioned electronic devices may include mobile phones and other electronic devices, i.e., various terminal devices used by users, also referred to as user devices, etc., without limitation. The aforementioned first application may be a memo application, photo gallery application, image enhancement application, video editing application, or creative application, document application, or instant messaging application installed on mobile phones and other electronic devices. This first application needs to call artificial intelligence tools, such as AI creation services or AI generation services, or AI generation tools, to provide users with editing or creation functions, referred to as creation functions. The aforementioned reference data may be the data referenced by the personalized content (such as the aforementioned first processing result) required by the user generated through the AI creation service or AI generation service, or the data referenced in association, also referred to as associated data.
[0009] Based on the above data processing method, this application enables mobile phones and other electronic devices to process the target content (i.e., the first display content mentioned above) according to the associated reference data (i.e., reference data or associated data) when they detect an instruction to edit or create the target content (i.e., the first display content mentioned above), thereby controlling the generation of personalized content that accurately meets the user's needs and improving the user experience.
[0010] It is understood that the aforementioned first processing instruction can be triggered by the user clicking on a processing control on the relevant interface, such as clicking on an AI generation tool or a writing assistant on the relevant interface of a memo application. In some embodiments, the aforementioned first processing instruction can also trigger the display of more processing controls. For example, clicking on an AI generation tool on the relevant interface of a memo application can trigger a function list interface on a mobile phone or other electronic device, which includes multiple editing functions such as a writing assistant, poetry and prose, AI drawing, and story creation. Clicking on any processing control corresponding to any editing function in the function list interface can trigger further processing instructions to process the aforementioned first displayed content, thereby obtaining the aforementioned first processing result.
[0011] In one possible implementation of the first aspect described above, the reference data is data stored in an electronic device. That is, the reference data can be local data from electronic devices such as mobile phones, including relevant data or descriptive information stored locally in such devices, without limitation.
[0012] In this way, when responding to user operations and calling AI generation tools to generate personalized content for users, mobile phones and other electronic devices can link local data, making the generated personalized content more in line with user needs and improving the user experience.
[0013] In one possible implementation of the first aspect above, the reference data includes one or more of the following: image data; video data; audio data; calendar data; map data; email data; text data; contact data; and SMS data.
[0014] In one possible implementation of the first aspect above, the reference data includes application data of the second application, wherein the application data of the second application includes data generated during the operation of the second application by the electronic device, and / or data accessed by the second application from the local storage space of the electronic device.
[0015] In other words, the aforementioned reference data can be application data provided by another application, and this other application, i.e., the second application, can be different from the application of the first application. For example, if the first application is a memo application, the second application can be a gallery application. Correspondingly, the aforementioned reference data can be photos or videos, or images or videos, provided by the gallery application. The photos or videos provided by the gallery application can be photos or videos taken by the user using the camera function of a mobile phone or other electronic device, or images or videos downloaded or saved by the user through a webpage or other applications, etc., without any restrictions.
[0016] In one possible implementation of the first aspect above, the second application includes one or more of the following applications: gallery application; recording application; calendar application; map application; email application; notepad application; text editing application; contact application; SMS application.
[0017] The aforementioned reference data can be image data and / or video data provided by gallery applications, audio data provided by recording applications, calendar data provided by calendar applications, map data provided by map applications, email data provided by email applications, text data provided by notepad or other text editing applications, contact data provided by contact applications, and SMS data provided by SMS applications, etc., without any restrictions.
[0018] In this way, when mobile phones and other electronic devices generate or create personalized content that meets user needs using correlated data, they can integrate data from multiple types of applications and combine correlated data from multiple dimensions to make the generated personalized content more in line with user needs. For example, integrating image or video data from photo library applications can allow the generated personalized content to incorporate data on dimensions such as the characteristics of the people, relationships between people, and environmental features that the user wants to include. Similarly, integrating calendar data from calendar applications can make the generated personalized content more closely match the user's schedule. Furthermore, integrating map data from map applications can make the generated personalized content more relevant to locations the user has visited or saved.
[0019] In one possible implementation of the first aspect above, when the reference data includes image data or video data, the information obtained from the reference data includes one or more of the following: character feature information of one or more people in the image data or video data; relationship information of multiple people in the image data or video data; and environmental feature information of the background of one or more people in the image data or video data.
[0020] In one possible implementation of the first aspect mentioned above, the character feature information includes one or more of the following: appearance features, action features, clothing features, age features, and gender features.
[0021] The physical characteristics mentioned above can include features related to a person's appearance or personal image, such as black hair, fair skin, dark skin, long hair, short hair, height, and weight. The action characteristics mentioned above can include activities such as playing soccer, riding a bicycle, playing in the sand, and flying a kite. The clothing characteristics mentioned above can include details about the person's clothing style (including color and coordination), accessories, etc., reflecting their fashion sense. The age characteristic mentioned above, as the name suggests, can include the person's age. The gender characteristic mentioned above, as the name suggests, can also include information such as gender, which will not be elaborated upon here.
[0022] The aforementioned relationship information may include relationships indicated by names added by users through photo library applications, such as names added by users in the voice of their children to photos in a photo library application, including father, mother, younger brother, grandfather, grandmother, maternal grandfather, maternal grandmother, etc. This relationship information may also include relationships determined from group photos or videos in photo library applications; there is no limitation on this.
[0023] The aforementioned environmental characteristics may include environmental features such as natural scenery and architectural style, or environmental features such as sports venues or sports fields with specific functions, without any restrictions.
[0024] For example, using image and / or video data provided by a photo library application as reference data, the information obtained based on the reference data in the first processing result described above can include character feature information, relationship information, and environmental feature information determined based on the aforementioned image and / or video data, and these feature information can be interconnected. For instance, the relationship information for "father" might correspond to the character feature information of being "tall and thin," the relationship information for "mother" might correspond to the character feature information of having "fair skin and long hair," and the relationship information for "father" might correspond to environmental feature information such as "riding a bicycle or playing soccer."
[0025] In one possible implementation of the first aspect above, processing the first displayed content based on reference data to generate a first processing result includes: obtaining at least one prompt word based on the reference data, and processing the first displayed content based on the at least one prompt word to generate the first processing result.
[0026] The aforementioned process of processing the first displayed content based on reference data includes first processing the reference data into prompts, i.e., prompt information. For example, a prompt generation model can be used to convert the reference data into prompts or prompt information. Then, based on the prompt information obtained from the above conversion, personalized content that can accurately match the user's needs is generated, i.e., the first processing result, such as an essay titled "My Family" generated by a writing assistant.
[0027] It is understandable that the above-mentioned transformation of reference data into prompt information can include extracting keywords or key phrases from text-based reference data to generate prompt information, or extracting image description information (i.e., a description of the reference data) from image-based reference data, and then processing it to generate prompt information. Furthermore, based on the prompt information generated above, further processing is performed to obtain personalized content that accurately matches the user's needs, i.e., the first processing result, such as an essay titled "My Family" generated by a writing assistant.
[0028] In some embodiments, when the reference data is image or video data, the description of the reference data may include descriptions of character features, relationships, and environmental features extracted from the reference data. For example, for an essay titled "My Family," the reference data may be image or video data related to characters such as "Dad," "Mom," and "Brother" provided by a photo library application. The description of the reference data may include feature data such as physical features and movement features extracted from characters such as "Dad," "Mom," and "Brother." Correspondingly, the first processing result generated based on the reference data, such as an essay titled "My Family," may include information obtained from the reference data, such as character feature information and relationship information of characters such as "Dad," "Mom," and "Brother." It may also include a family photo redrawn based on this information, or individual images of characters such as "Dad," etc., without limitation.
[0029] In one possible implementation of the first aspect above, obtaining at least one prompt word based on reference data includes: determining at least one control parameter based on the reference data, wherein the control parameter includes the reference data or a description of the reference data; and performing a transformation process on the at least one control parameter to obtain at least one prompt word.
[0030] In some embodiments, during the process of obtaining prompt words or information based on the aforementioned reference data, control parameters can be determined by first extracting a description of the reference data from the reference data, or by using the reference data itself to determine the control parameters, and then inputting them into the prompt generation model to generate prompt words or information.
[0031] In one possible implementation of the first aspect above, the electronic device includes a first model, and processes the first displayed content based on at least one prompt word to generate a first processing result, including: invoking the first model, and inputting the first displayed content and at least one prompt word into the first model for model inference to generate the first processing result.
[0032] That is, the first model mentioned above can perform model reasoning based on the first displayed content and at least one prompt word to generate the first processing result.
[0033] For example, the first model mentioned above can be a large model developed based on AIGC technology, namely an artificial intelligence model or AI model. This model can provide various AI generation capabilities and can perform model inference based on the first display content of the input AI model and the prompt words obtained from the reference number, etc., to generate personalized content that meets the user's needs, namely the first processing result mentioned above.
[0034] It is understood that the above-mentioned processing of the first display content based on reference data, if it is necessary to control the first processing result generated accordingly, including the information obtained based on the reference data, can extract data related to character characteristics, data related to character relationships, and data related to environmental characteristics based on the above-mentioned reference data, and convert them into prompt information input to the above-mentioned first model (large model) or related components based on the large model.
[0035] In one possible implementation of the first aspect described above, the at least one control parameter further includes a parameter determined by the user's parameter tuning operation on the at least one parameter tuning control.
[0036] In some embodiments, the electronic device may also display parameter tuning controls that allow the user to set relevant control parameters that can affect the first processing result. Correspondingly, based on the user's tuning operations on these parameter tuning controls, the aforementioned at least one control parameter may include the result of these tuning operations.
[0037] It is understood that the aforementioned first processing result is related to constraints, which are determined by the first display content and at least one control parameter, and are used to constrain the first processing result output by the aforementioned first model. Based on this, the constraints can be determined, on the one hand, by the first display content input to the first model, and on the other hand, by at least one control parameter capable of determining the prompt information of the input model.
[0038] In one possible implementation of the first aspect above, when the first processing result includes text, at least one control parameter includes one or more of the following: the number of words used to limit the length of the text content; the tone used to limit the descriptive style of the text content; and the target audience category used to limit the text content.
[0039] For example, taking the essay generated by the writing assistant as the first processing result mentioned above, at least one of the above control parameters, such as the control parameters "5000", "formal", "university", etc., set by the user through the relevant parameter adjustment controls, can be included. The parameter adjustment controls corresponding to the above applicable user categories may include, for example, selection controls corresponding to the word "stage" such as "primary school", "junior high school", "high school", "university", and "work / other", etc., without limitation.
[0040] In one possible implementation of the first aspect above, when the first processing result includes an image or video, at least one control parameter includes one or more of the following: a style parameter for defining the style of the image or video; a size parameter and / or material parameter for defining the clothing of the main subject in the image or video; a shot size parameter and angle parameter for defining the prominence of the main subject in the image or video; a parameter for defining the age of the person in the image or video; and a parameter for defining the location of the person in the image or video.
[0041] For example, the style parameters mentioned above can include parameters related to changing the style of a person in a user-specified image within the relevant interface of a memo application. The size parameters mentioned above can include personal size parameters such as the user's height, chest circumference, waist circumference, and hip circumference. The material parameters mentioned above can include parameters related to clothing material, such as "pure cotton." The shot parameters mentioned above can include parameters related to shot type, such as "close-up" and "long shot." The angle parameters mentioned above can include parameters related to the person's posture, such as "side view," "front view," and "back view." The age parameters mentioned above can include parameters used to define the age of a person in an image or video, such as "young male" or "teenage male." The location parameters mentioned above can include parameters such as "track and field," "swimming pool," or "gym."
[0042] In one possible implementation of the first aspect above, before processing the first displayed content according to the reference data, the method further includes: detecting a first processing instruction for the first displayed content and displaying a second interface, wherein the second interface includes an associated control, the associated control being used to trigger the acquisition of reference data required to generate the first processing result.
[0043] Before processing the first displayed content based on the reference data, the method further includes: detecting that the associated control is in an open state and obtaining the reference data.
[0044] In some embodiments, a user-instructed first processing command to process the first displayed content can trigger an electronic device such as a mobile phone to display an interface including associated controls, such as the second interface described above. On this second interface, the user can control the on / off state of the associated controls to control whether associated data is acquired, thereby affecting the data processing result, such as the first processing result described above.
[0045] In some embodiments, the associated control on the second interface may be a switch control for associating with a gallery application, calendar application, or map application. When the user clicks the switch control to turn it on, the mobile phone or other electronic device can associate relevant data (i.e., local data) stored locally when it calls a generation model such as the writing model to generate the content required by the user. This includes application data generated during the operation of the gallery application, calendar application, or map application, as well as other local data accessible by each application.
[0046] In one possible implementation of the first aspect above, detecting that the associated control is in an open state and obtaining reference data includes: displaying a selection control corresponding to at least one category of data related to the reference data, wherein the at least one category of data includes a first category of data and the first category of data corresponds to a first selection control; detecting a first operation of the user selecting the first selection control, and obtaining the first category of data as reference data.
[0047] For example, the first category of data mentioned above may include various types of data related to "My Family" that can be displayed on the writing assistant interface below, such as data with "Dad", "Mom" and "Brother" as category labels.
[0048] In one possible implementation of the first aspect described above, the second interface further includes at least one parameter tuning control, and the method further includes: detecting a user's parameter tuning operation on at least one parameter tuning control, processing the first display content according to reference data and parameters determined by the parameter tuning operation, and generating a first processing result.
[0049] It is understood that the user's parameter tuning operations on at least one parameter tuning control detected above can correspond to the parameters determined by the user's parameter tuning operations on at least one parameter tuning control. That is, the user's parameter tuning operations on at least one parameter tuning control can change the parameter content of the at least one control parameter, thereby adjusting the content of the prompt information generated by the prompt generation model, so as to affect the content of the first processing result output by the corresponding input AI model, i.e., the first model after model inference.
[0050] It is understood that the second interface mentioned above may be the interface provided by the AI generation service or AI generation tool called by the first application, or it may be the interface of the third application called by the first application, such as the interface of the AI application called by the memo application, etc. There are no restrictions here.
[0051] In one possible implementation of the first aspect described above, the second interface further includes a generation control indicating the start of generating a processing result, and the method further includes: detecting a second operation by the user clicking the generation control, and displaying a third interface, wherein the third interface includes a first processing result obtained by processing the first displayed content according to reference data.
[0052] For example, the aforementioned generation control may include the "Start Generation" control or the Regenerate control described below. The aforementioned third interface may be an interface that includes the generation results, such as the various generation result interfaces exemplified below, which will not be elaborated upon here.
[0053] In one possible implementation of the first aspect described above, the third interface further includes an application control that applies the first processing result to the first interface, and displays the first processing result on the first interface, including: detecting a third operation by the user clicking the application control, and displaying the first processing result on the first interface.
[0054] For example, the application control described above may include the application control shown in the example below, which is displayed on the result generation interface and can apply the generated result (such as the first processing result described above) to the first interface displaying the first display content.
[0055] In one possible implementation of the first aspect above, when the first application is a memo application and the first displayed content includes text or an image, a first processing instruction for the first displayed content is detected, including any one of the following: the first processing instruction is detected based on a fourth operation where the user long-presses text in the first interface and selects the first displayed content; or the first processing instruction is detected based on a fifth operation where the user long-presses the first displayed content.
[0056] When the first application is any one of a gallery application, an image enhancement application, or a creative application, and the first displayed content includes an image or video, a first processing instruction for the first displayed content is detected, including any one of the following: the first processing instruction is detected based on a sixth operation of the user long-pressing the first displayed content; the first processing instruction is detected based on a seventh operation of the user clicking on a processing control on the first interface, and the processing control on the first interface includes a control that indicates processing the first displayed content on the first interface.
[0057] In one possible implementation of the first aspect described above, the detection of a first processing instruction on the first displayed content may include an instruction response process that triggers the display of a multi-level control menu. For example, before detecting the first processing instruction, an electronic device such as a mobile phone may display a first-level control menu upon detecting a second processing instruction from the user on the first displayed content. This first-level control menu may, for example, include a function option box for AI generation tools, as described below. Furthermore, the electronic device may detect the first processing instruction based on the user's selection operation on the AI generation tool within the first-level control menu and display a second-level control menu. This second-level control menu may include multiple processing controls within the category of AI generation tools, such as writing assistants, poetry and prose, AI painting, and story creation.
[0058] It is understood that the processing controls included in the first-level control menu or the second-level control menu mentioned above may include one or more of the following controls:
[0059] A control used to trigger text expansion functionality, such as a writing assistant. This text expansion functionality generates first text content based on first displayed content and constraints. This first text content may include text from the first displayed content or keywords extracted from the first displayed content.
[0060] A control used to trigger the function of matching poems and songs, wherein the function of matching poems and songs is used to obtain the second text content according to the constraints, wherein the form of the second text content includes poems and songs.
[0061] Controls used to trigger drawing functions, such as AI drawing. These drawing functions generate a first image based on constraints. The first image may include a person with certain physical characteristics, such as a "tall and thin" father or a "fair-skinned, long-haired" mother; it may also include natural or cultural landmarks with certain scenic features, or combinations of people and / or animals with certain related characteristics, such as a "girl holding a dragon," etc., without limitation.
[0062] This control triggers story creation functions, such as story creation or story inspiration. The story creation function generates third-party text content based on constraints; the literary genre of the third-party text content is a story.
[0063] A control used to trigger an image processing function. This image processing function is used to process a first displayed content into a second image based on constraints, the second image being the same as the foreground subject in the first displayed content.
[0064] A control used to trigger functions that change the style of an image or video.
[0065] A control used to trigger functions that change the clothing and accessories of people in images or videos.
[0066] A control used to trigger the ability to generate a poster. This poster generation ability is used to generate a third image, which is a poster with promotional effects, based on the first displayed content and constraints.
[0067] A control used to trigger functions that repair the clarity of images or videos.
[0068] The control is used to trigger the partial modification function, which is used to modify or redraw a part of the first displayed content according to the first displayed content and constraints.
[0069] This control is used to trigger the image expansion function, which expands the size of the background area in the first displayed content according to constraints.
[0070] This control is used to trigger the background blurring function, which blurs the background of the first displayed content according to constraints.
[0071] Based on the data processing method provided by the first aspect and related implementations mentioned above, users can set control parameters for specified target content such as text, images, audio, and video on the relevant interface of the target application displayed on mobile devices such as smartphones. These control parameters can also be associated with locally stored related data and corresponding generated descriptive information. Ultimately, this allows for the control of generating personalized content that accurately meets the user's needs, thus improving the user experience. Furthermore, the subsequent use of the generated results can be achieved by clicking on relevant controls to directly apply the results to the application interface the user is currently using, i.e., the first interface of the aforementioned first application, such as the relevant interface of a memo application, without requiring complex switching operations from the user, which also improves the user experience.
[0072] In a second aspect, this application provides an electronic device, including: one or more processors; one or more memories; the one or more memories storing one or more programs, which, when executed by one or more processors, cause the electronic device to perform the data processing methods provided in the first aspect and various possible implementations of the first aspect.
[0073] Thirdly, this application provides a computer-readable medium storing instructions that provide a data processing method in accordance with the first aspect and various possible implementations thereof.
[0074] Fourthly, this application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the data processing method provided in the first aspect and various possible implementations of the first aspect.
[0075] The beneficial effects of the second to fourth aspects mentioned above can be referred to the relevant descriptions in the first aspect and various possible implementations of the first aspect, which will not be repeated here. Attached Figure Description
[0076] Figure 1 shows a schematic diagram of a data processing scenario that provides text expansion functionality through a specific application interface.
[0077] Figure 2a shows a schematic diagram of the memo interface involved in generating content required by the user through the target application calling generation model provided in this embodiment of the application.
[0078] Figure 2b shows a schematic diagram of the text editing interface involved in the generation of expanded content by calling the target application generation model according to the embodiment of this application.
[0079] Figure 2c shows a schematic diagram of the writing assistant interface involved in the target application calling the generation model to generate expanded content according to the embodiment of this application.
[0080] Figure 2d shows a schematic diagram of the generation result interface involved in the target application calling the generation model to generate expanded content according to the embodiment of this application.
[0081] Figure 2e shows a schematic diagram of a memo interface including a first processing result provided in an embodiment of this application.
[0082] Figure 3 shows a schematic diagram of the implementation process of a data processing method provided in Embodiment 1 of this application.
[0083] Figure 4a shows a schematic diagram of an application scenario where a memo application call generation model is integrated with associated data to generate expanded content, as provided in Embodiment 2 of this application.
[0084] Figure 4b shows a schematic diagram of a text editing interface provided in the application scenario shown in Figure 4a.
[0085] Figure 4c shows a schematic diagram of a writing assistant interface provided in the application scenario shown in Figure 4a.
[0086] Figure 4d shows a schematic diagram of a result generation interface provided in the application scenario shown in Figure 4a.
[0087] Figure 4e shows a schematic diagram of a memo interface that includes a second processing result in the application scenario shown in Figure 4a.
[0088] Figure 4f shows a schematic diagram of a schedule interface that provides schedule data as associated data.
[0089] Figure 4g shows a schematic diagram of a map interface that provides map data as associated data.
[0090] Figure 5 shows a schematic diagram of the implementation process of a data processing method provided in Embodiment 2 of this application.
[0091] Figure 6a shows a schematic diagram of an application scenario where an image of a memo interface needs to be processed into a virtual avatar, according to Embodiment 3 of this application.
[0092] Figure 6b shows a schematic diagram of a memo interface that includes at least one editing function, provided in the application scenario shown in Figure 6a.
[0093] Figure 6c shows a schematic diagram of an image editing interface provided in the application scenario shown in Figure 6a.
[0094] Figure 6d shows another image editing interface provided in the application scenario shown in Figure 6a.
[0095] Figure 6e shows a schematic diagram of a result generation interface provided in the application scenario shown in Figure 6a.
[0096] Figure 6f shows a schematic diagram of a memo interface that includes a third processing result in the application scenario shown in Figure 6a.
[0097] Figure 7 shows a schematic diagram of the implementation process of a data processing method provided in Embodiment 3 of this application.
[0098] Figure 8a shows a schematic diagram of an application scenario where the clothing of a person in a photo on a gallery interface needs to be modified, according to Embodiment 4 of this application.
[0099] Figure 8b shows a schematic diagram of an AI-generated interface provided in the application scenario shown in Figure 8a.
[0100] Figure 8c shows a schematic diagram of a result generation interface provided in the application scenario shown in Figure 8a.
[0101] Figure 8d shows a schematic diagram of a gallery interface that includes the fourth processing result in the application scenario shown in Figure 8a.
[0102] Figure 9 shows a schematic diagram of the implementation process of a data processing method provided in Embodiment 4 of this application.
[0103] Figure 10a shows a schematic diagram of an application scenario where storyboard creation is required to generate creative storyboard images or storyboard videos, as provided in Embodiment 5 of this application.
[0104] Figure 10b shows a schematic diagram of an AI-generated interface provided in the application scenario shown in Figure 10a.
[0105] Figure 10c shows a schematic diagram of a result generation interface provided in the application scenario shown in Figure 10a.
[0106] Figure 10d shows a schematic diagram of a gallery interface that uses the generated results in the application scenario shown in Figure 10a.
[0107] Figure 11 shows a schematic diagram of the implementation process of a data processing method provided in Embodiment 5 of this application.
[0108] Figure 12a shows a schematic diagram of the software structure of an operating system suitable for mobile phones and other electronic devices provided in an embodiment of this application.
[0109] Figure 12b shows a schematic diagram of the structure of a generative model provided in an embodiment of this application.
[0110] Figure 12c shows a schematic diagram of the structure of an AI subsystem provided in an embodiment of this application.
[0111] Figure 12d shows a schematic diagram illustrating the implementation principle of generating expanded content by incorporating associated data, as provided in an embodiment of this application.
[0112] Figure 13 shows a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0113] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0114] Prompts are a form of input guidance, which can be understood as the language used to communicate with AI models (such as the large models mentioned above). They describe the features of the text, images, audio, video, etc., that the AI model wants to generate. The accuracy and precision of the prompts largely determine whether the generated content meets the user's expectations.
[0115] It should also be stated that the steps in the methods and processes in this application are numbered for ease of reference, not to limit the order of steps. If there is an order between the steps, the textual description shall prevail.
[0116] It is understood that the terminal device in the embodiments of this application may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. Terminal devices may include mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and so on.
[0117] As mentioned earlier, current application interfaces based on large models that provide functions such as writing assistants and AI drawing require users to input relatively specialized prompts as model input, and cannot be linked to application data from other local applications. This results in a poor user experience because the content generated by these large model-based applications (hereinafter referred to as AI applications) differs significantly from the user's expectations.
[0118] Figure 1 illustrates a data processing scenario that provides text expansion functionality through a specific application interface.
[0119] Taking mobile phone 100 as an example, as shown in Figure 1, when mobile phone 100 runs a specific application based on a large model (hereinafter referred to as AI application), it can display the AI application interface 101. This interface 101 can include a "writing assistant" function 102 and functions such as "AI drawing" and "movie recommendation". Users can select the "text expansion" function 102 on this interface 101, and correspondingly, mobile phone 100 can display the writing assistant interface 103 shown in Figure 1 based on the text expansion function provided by the large model. Then, users can enter detailed writing requirements in the input box 103a of the writing assistant interface 103, or paste writing materials or original text content that needs to be expanded. Alternatively, users can also paste writing materials or original text content that needs to be expanded in the input box 101a of the AI application interface 101, and specify detailed writing requirements, such as word count, tone, and age group of the target audience, etc.
[0120] Referring again to Figure 1, after the user selects the "Writing Assistant" function 102 and enters detailed writing requirements, the writing assistant interface 103 displayed on the mobile phone 100 can then display the corresponding generated content 103b, such as the essay or expanded content. Whether this content 103b can meet the user's needs often depends on whether the writing requirements entered by the user in the input boxes 103a or 101a can accurately describe the user's ideas.
[0121] However, when users input detailed writing requirements, due to a lack of professional knowledge in the relevant field, they may not be able to accurately describe the professional prompts they want to expand on. This will result in the generated content or expanded content being far from what the user expects, leading to a poor user experience.
[0122] To address the aforementioned issues, this application provides a data processing method. By offering users an interface on which editing or creation functions based on a large model are implemented, users can select and associate local data. This allows users to choose to associate local data, ensuring that the data processing results indicated on the interface better meet user needs. The resulting personalized content thus better aligns with user expectations, improving the user experience. Furthermore, the interface can provide at least one parameter tuning control related to the generated results, guiding users to adjust parameters to control the processing outcome. This controls the model inference process on the input content, ultimately generating or creating content such as the aforementioned expanded content.
[0123] The control parameters that users configure for target content on the corresponding function panels can be transformed into corresponding prompt information that accurately describes the user's relevant needs using a pre-trained prompt generation model. This information serves as input to the corresponding generation model, ensuring that the final processing result better matches the user's requirements. In this way, different parameter tuning controls can prompt users to set control parameters as input prompts, thereby generating accurate prompt information to better match the processing result with the user's needs and improve the usability of the generated personalized content.
[0124] The interfaces related to the various types of generation capabilities or creation capabilities based on the aforementioned large model can be AI generation interfaces displayed by the target application running on mobile devices or other electronic devices, which call the preset generation model in the system. They can also be interfaces related to the AI applications developed based on the aforementioned large model; no limitation is made here. The aforementioned generation models can include generation models with different generation capabilities, and the parameters dependent on which each generation model generates the content required by the user can be adjusted based on the user's settings of the corresponding parameter tuning controls. Therefore, in some embodiments, the aforementioned generation models can also be called controllable generation models or adjustable parameter generation models, etc., without limitation. Simultaneously, based on the data processing method provided in this application, each generation model can pre-register the required large model capabilities. For example, the generation model that a memo application can call can include components built from the large model capability that provides text expansion functionality to generate expanded content. Furthermore, the corresponding generation model can integrate a parameter tuning module to provide the target application with parameter tuning controls on the corresponding function panel for setting relevant parameters. For example, for the writing assistant generation model (referred to as the writing model), parameter tuning controls can be set to control the word count, tone, and target audience of the essay or expanded content.
[0125] In some embodiments, the generated model may be a system service or system component provided by the operating system of an electronic device such as a mobile phone, such as an AI creation service or an AI generation service, an AI model component, etc., and there are no restrictions on this.
[0126] The target applications mentioned above can include various applications installed on mobile devices such as mobile phones, such as system applications such as memo applications and gallery applications, as well as third-party applications such as image enhancement applications, video editing applications or creative applications, document applications or instant messaging applications.
[0127] It is understandable that when processing target content of the same target type (including text, images, audio, video, etc.), the application can call a generative model that supports the corresponding type of function to implement the processing. For example, a memo application can call a generative model with text expansion capabilities to implement functions such as a writing assistant (e.g., the first type of function below) for the text selected by the user (hereinafter referred to as the target text); a memo application can also call a generative model with image processing capabilities to implement image processing functions (e.g., the second type of function below) for the target image. When processing target content of the same modality, different target applications can call a generative model that supports the same type of function to implement the processing. For example, a memo application can call a generative model with image processing capabilities to implement image processing functions, and a gallery application can also call a generative model with image processing capabilities to implement image processing functions. No restrictions are imposed here.
[0128] The aforementioned large models may include large pre-trained language models pre-deployed in the operating system (OS) of mobile devices such as mobile phones, or large pre-trained language models loaded from the corresponding cloud or server when artificial intelligence (AI) applications installed on mobile devices such as mobile phones are running. The aforementioned AI applications may include, for example, Wenxin Yiyan™, Tongyi Qianwen™, etc., without limitation.
[0129] As an example, Figures 2a to 2e illustrate relevant interface diagrams for calling the generative model to generate content required by the user on the target application, according to embodiments of this application.
[0130] Taking the memo application as an example, as shown in Figure 2a, on the memo interface 210 displayed on the mobile phone 100, the user can select the target text. The mobile phone 100 can display the available functions through the function option box 211. As an example, the function option box 211 can display various function options that can call up the large model (or artificial intelligence model), such as AI generation tool 212, writing assistant 213, and poetry 214. Among them, AI generation tool 212 corresponds to different types of functions in the opened function interface, including writing assistant 213, poetry 214, etc., which can be provided by generation models with different types of generation capabilities, or by the same generation model with a certain type of generation capability, without restriction. Referring to Figure 2a again, the function option box 211 displayed on the mobile phone 100 can also display options such as "copy" function 215 and "cut" function 216, providing the user with the functions of copying and cutting the selected text, respectively. These functions do not have to be basic functions implemented based on the large model, without restriction.
[0131] As shown in Figure 2b, the mobile phone 100 can respond to the user's operation of clicking the AI generation tool 212 in the function option box 211 of the aforementioned memo interface 210, and display a text editing interface 220 with various types of functions. For example, the text editing interface 220 may include functions such as writing assistant 221, poetry and prose 222, AI drawing 223, and story creation 224. Each function can be implemented by calling a generation model with corresponding generation capabilities, or by calling a generation model with multiple generation capabilities; no restrictions are placed here. Among them, writing assistant 221 can correspond to writing assistant 213 provided in the function option box 211 of the aforementioned memo interface 210, and writing assistant 213 can serve as a shortcut function option for the writing assistant 221 function in the aforementioned function option box 211. Similarly, poetry and prose 222 can correspond to poetry and prose 214 provided in the function option box 211 of the aforementioned memo interface 210, and so on, which will not be elaborated here.
[0132] In some embodiments of this application, the editing functions displayed in the aforementioned function option box 211 can serve as editing functions provided by the first-level control menu; the editing functions in the text editing interface 220 corresponding to the AI generation tool 212 displayed when clicking on the function option box 211 can serve as editing functions provided by the second-level control menu. Thus, target applications such as memo apps, based on the corresponding capabilities provided by the large model, can display two or more levels of control menus to the user, providing more and richer editing functions to meet the user's needs for different editing functions and improve the user experience.
[0133] Referring again to Figure 2c, when the mobile phone 100 detects that the user clicks on the writing assistant 221 on the text editing interface 220, the writing assistant interface 230 can be displayed accordingly. This writing assistant interface 230 provides several parameter adjustment controls, including "300," "800," "1000," "2000," "5000," "10000," and "custom" for "word count," "friendly," "professional," "formal," "humorous," "casual," and "direct" for "tone," and "applicable audience categories" such as "primary school," "junior high school," "high school," "university," and "work / other," etc., which are not limited here. The "stage" can refer to the applicable audience category corresponding to the generated writing content. In other embodiments, this applicable audience category may also include other classification methods, such as applicable audience categories based on different occasions or usage scenarios, which are also not limited here.
[0134] For example, if a user performs parameter adjustments on the writing assistant interface 230, such as selecting "5000," "Formal," or "University," the mobile phone 100 can display "Word Count: 5000," "Tone: Formal," and "Stage: University" as prompts in the description information input box 232 of the writing assistant interface 230. Furthermore, if the mobile phone 100 detects that the user clicks the "Start Generation" control 231 on the writing assistant interface 230, it can combine the parameter adjustments performed by the user on the controls for dimensions such as "Word Count," "Tone," and "Stage," using the target text as input to the writing model, and ultimately generate the content required by the user, i.e., an essay or expanded content that meets the user's expectations.
[0135] In this way, users do not need writing skills or professional knowledge related to writing models; they can simply click on the various parameter adjustment controls or set control parameters to input accurate prompt information and control the writing model and other generation models to generate the desired content. That is, the data processing method provided in this application facilitates user operation and achieves high accuracy in the data processing results.
[0136] Referring again to Figure 2d, the generation result interface 240 displayed by the mobile phone 100 may include the generated extended content 240a and operation controls corresponding to the processing result, such as application control 241, copy control 242, regenerate control 243, and share control 244. The application control 241 can trigger the application of the generated extended content 240a to the target application that calls the corresponding generation model, such as the relevant interface of a memo application. For example, if the mobile phone 100 detects that the user clicks the application control 241 on the generation result interface 240, the generated extended content 240a can be automatically filled into the relevant interface of the memo application. At this time, the mobile phone 100 may display, for example, the memo interface 250 shown in Figure 2e, which displays the processing result shown in the generation result interface 240 of Figure 2d, i.e., the extended content 240a. In other words, compared to the memo interface 210 shown in Figure 2a, the memo interface 250 shown in Figure 2e can display the extended content 240a directly applied to the memo interface.
[0137] In other embodiments, users can also click the copy control 242 on the generated result interface 240 to copy the generated expanded content 240a and paste it into the interface of other applications, such as the chat interface of an instant messaging application. Users can also click the regenerate control 243 on the generated result interface 240 to instruct the expanded content 240a to be regenerated. Users can also click the share control 244 on the generated result interface 240 to share the generated expanded content 240a with friends. In other embodiments, the generated result interface 240 may also include other controls for performing related operations on the displayed expanded content 240a, such as pagination controls "Next Page," "Previous Page," etc., which are not limited here.
[0138] Furthermore, the data processing method provided in this application, while offering the aforementioned parameter adjustment controls for users to perform parameter adjustment operations, can also set associated controls that can be linked to local data within the provided parameter adjustment controls. In some embodiments, the associated control can be, for example, a switch control for linking to a gallery application, calendar application, or map application. When the user clicks the switch control to turn it on, the mobile phone or other electronic device can link relevant locally stored data (i.e., local data) when calling a generation model such as the aforementioned writing model to generate the content required by the user. This includes application data generated during the operation of the aforementioned gallery application, calendar application, or map application, as well as other local data accessible by each application, etc., without limitation. If it is not necessary to link local data, the user can also turn the aforementioned switch control off, which will not be elaborated here.
[0139] The application data of the aforementioned image library application may include, for example, image or video data accessible by the application. This data, as reference data, can provide related materials such as people, scenery, and events for the content generated by the writing model. For example, the writing example below, titled "My Family," can be generated by associating the image library data of the image library application with the user's expectations. For example, the content of the essay may include "Dad is tall and thin," often taking us to "ride bicycles and play soccer," "Mom has fair skin and long hair," "My younger brother paints his face in various colors," and likes to make "creative" shapes, etc. Similarly, the schedule data of the aforementioned calendar application can provide related materials such as time, events, and plans; the map data of the map application can provide location-based map check-in data (including time, location, and other information). In some embodiments, the aforementioned associated application data may also include application data received, acquired, or collected during the operation of applications such as email applications, recording applications (e.g., recorders), contacts, and SMS, etc., without limitation.
[0140] In this way, when mobile phones and other electronic devices generate or create personalized content that meets user needs using correlated data, they can integrate data from multiple types of applications and combine correlated data from multiple dimensions to make the generated personalized content more in line with user needs. For example, integrating image or video data from photo library applications can allow the generated personalized content to incorporate data on dimensions such as the characteristics of the people, relationships between people, and environmental features that the user wants to include. Similarly, integrating calendar data from calendar applications can make the generated personalized content more closely match the user's schedule. Furthermore, integrating map data from map applications can make the generated personalized content more relevant to locations the user has visited or saved.
[0141] In the embodiments of this application, since the above reference data can provide related materials for generating corresponding processing results for writing models, etc., it can also be referred to as related data below.
[0142] It is understood that the target application running on mobile phone 100 can also include gallery applications, content creation applications, image enhancement applications, or other similar applications. Correspondingly, based on the data processing method provided in this application, mobile phone 100 running the target application can call generation models with different types of generation capabilities. These different types of generation capabilities can correspond to the capabilities provided by each generation model to edit or trim data such as text, images, audio, or video, or to perform reasoning and creation to generate relevant personalized content that meets user needs. For example, based on the data processing method provided in this application, mobile phone 100 can call the generative model with image processing function provided by the large model to perform image enhancement processing or generate virtual images on user-specified images; as another example, mobile phone 100 can call the generative model with certain creative capabilities, such as the advertisement generation model, wallpaper generation model, story creation model, and planning document generation model, to generate creative advertisement storyboards, wallpapers, stories, or planning documents on user-specified text, images, audio, or video; as yet another example, mobile phone 100 can call the video editing model with video editing capabilities to edit user-specified images and / or videos, etc., without any restrictions.
[0143] In this way, users can quickly call the generative model to generate the text, images, audio, or video content they need from the interface of the target application, such as a memo app running on mobile phones and other electronic devices. Users can also select the corresponding prompts on the relevant interface of the target application to adjust the parameters of the generative model, thereby making the generated content better meet the user's creative needs.
[0144] The following examples, using different application scenarios, detail the specific implementation process of the data processing method provided in this application.
[0145] The following section will first describe the application scenario of calling a generative model with text expansion function to generate expanded content based on the examples of the memo application in Figures 2a to 2e above, and then introduce the specific implementation process of a data processing method provided in this application in conjunction with Embodiment 1.
[0146] Example 1
[0147] This application uses a memo application as an example, illustrating the specific implementation process of the data processing method provided in this application embodiment based on the application scenario shown in Figures 2a to 2e above, where the memo calls a generative model with text expansion capabilities to generate expanded content. In this application scenario, the target application is a memo application. The memo application can provide functions such as a writing assistant based on a generative model with text expansion capabilities.
[0148] Figure 3 is a schematic diagram of the implementation flow of a data processing method according to an embodiment of this application.
[0149] As shown in Figure 3, in this embodiment of the application, the executing entity for each step in the implementation process can be a mobile phone 100. In other embodiments, the executing entity for each step in the implementation process shown in Figure 3 can also be other electronic devices, such as tablet computers, laptop computers, etc., and there is no limitation here.
[0150] Specifically, the implementation process may include the following steps:
[0151] S301: In response to the user opening the Notes app, the first screen of the Notes app is displayed.
[0152] For example, mobile phone 100 can respond to the user's action of clicking the icon of the memo application to run the memo application, and display the memo interface provided by the memo application, which allows the user to input multimodal data such as text, pictures, audio or video, such as the memo interface 210 shown in Figure 2a above. In order to distinguish it from the interface in the following embodiments, this memo interface can be referred to as the first interface.
[0153] In some embodiments, the mobile phone 100 may also respond to the user's click on the service card corresponding to the Notes application and display the first interface mentioned above. In other embodiments, the mobile phone 100 may first respond to the user's click on the Notes application icon and display the main interface of the Notes application, which can display all notes or memos. The mobile phone 100 may then respond to the user's click on one of the notes or memos on the main interface and display the Notes interface containing the content of that note, i.e., the first interface mentioned above. No limitations are imposed here.
[0154] S302: Detecting a user's selection of target text on the first interface, display at least one editing function on the first interface.
[0155] For example, the mobile phone 100 can detect when a user long-presses text in the current note on the memo interface, selecting part or all of the text as the target text, and then displays one or more editing functions that can be used to edit the selected target text. Here, the content of the current note displayed on the memo interface (i.e., the first interface mentioned above) is text, and the content selected by the user is the target text. In other embodiments, the content displayed on the first interface may also be images, audio, or video, and the content selected by the user may also be a target image, target audio, or target video, etc., without limitation. The editing functions displayed on the first interface may differ for different types of note content.
[0156] It is understood that the above-mentioned at least one editing function corresponding to the target text may include the first type of function implemented based on a large model or an artificial intelligence model. For example, it may include the various editing functions included in the AI generation tool 212 in the function option box 211 shown in Figure 2a, such as the writing assistant 221, poetry and prose 222, AI drawing 223, and story creation 224 shown in the text editing interface 220 shown in Figure 2b. In addition, the above-mentioned at least one editing function may also include some editing functions that are not implemented based on an AI model, such as the "copy" function 215 and the "cut" function 216 in the function option box 211 shown in Figure 2a.
[0157] Referring to Figure 2a above, the function option box 211 displayed by the mobile phone 100 can display either a first category of functions including multiple editing functions, such as the AI generation tool 212, or editing functions included in the first category of functions, such as the writing assistant 213 and poetry / song 213. The writing assistant 213 and poetry / song 213 may include editing functions randomly selected from the first category of functions or recommended based on user frequency. In other embodiments, the function option box 211 displayed by the mobile phone 100 may also simultaneously display the AI generation tool 212 and other editing functions belonging to the AI generation tool 212, such as AI drawing 223 or story creation 224, etc., without limitation.
[0158] It is understood that if the user clicks on a first-category function among at least one of the displayed editing functions, such as clicking on AI generation tool 212, then the mobile phone 100 can execute the response process of steps S303 to S306 below. If the user clicks on a recommended editing function belonging to the first-category function among at least one of the displayed editing functions, such as writing assistant 213, poetry and prose 213, etc., then the mobile phone 100 can execute the response process of steps S305 to S306 below. For specific execution processes, please refer to the relevant descriptions in the corresponding steps below, which will not be repeated here.
[0159] S303: The user has been detected to have selected at least one of the first type of editing functions.
[0160] For example, the first type of function among the above-mentioned at least one editing function may include, for example, an AI generation tool 212 with multiple editing functions. The mobile phone 100 can detect the user's click on the first type of function on the first interface displaying the at least one editing function. For example, this operation may include the user clicking on the AI generation tool 212 in the function option box 211 shown in FIG. 2a. Correspondingly, the mobile phone 100 can continue to execute the following step S304 in response to the user's selection of the first type of function among the at least one editing function. It is understood that the first type of function may include one or more editing functions that process data corresponding to a certain modality. Data of different modalities may include, for example, text, images, audio, and video data, etc., and are not limited here.
[0161] In some embodiments, the mobile phone 100 may also detect when the user selects at least one editing function, such as copy or cut. In this case, the mobile phone 100 can use the relevant capabilities of the memo application itself to perform the process of copying or cutting the text selected by the user in response to the user's operation.
[0162] S304: Displays the function list interface for the first category of functions.
[0163] For example, in response to the user's selection of at least one editing function of the first category, the mobile phone 100 can display a function list interface including all editing functions of the first category. As an example, referring to the text editing interface 220 shown in Figure 2b above, the text editing interface 220 can display various editing functions such as writing assistant 221, poetry and prose 222, AI drawing 223, and story creation 224.
[0164] The writing assistant 221 can be a control used to trigger a text expansion function. This text expansion function generates the user's desired essay or expanded content (denoted as the first text content) based on a first displayed content, such as the target text displayed on the first interface, and constraints formed by the first displayed content and control parameters set by the user through the aforementioned parameter-adjusting controls. In some embodiments, the first text content generated by the writing assistant 221 may include text from the first displayed content, or keywords extracted based on the first displayed content, etc., without limitation.
[0165] The "Poetry and Prose" control 222 can be a control with the function of matching poetry and prose. This function is used to match the corresponding poetry and prose (denoted as the second text content) based on the above constraints and the first displayed content, such as the target text displayed on the first interface. It can be understood that the poetry and prose refers to the literary form of the article, such as ancient poetry, poems, Chu Ci, Han Fu, etc., without restriction.
[0166] AI drawing 223 can be a control used to trigger the drawing function. This drawing function generates an image that meets the user's expectations, denoted as the first image, based on the aforementioned constraints and the first displayed content, such as the target text displayed on the first interface. The first image may include a person with certain physical characteristics, such as a "tall and thin" father or a "fair-skinned, long-haired" mother; it may also include natural or cultural landmarks with certain scenic features, or combinations of people and / or animals with certain related characteristics, such as a "girl holding a dragon," etc., without limitation.
[0167] The story creation 224 can be a control used to trigger the story creation function. This story creation function generates a story that meets the user's expectations based on the aforementioned constraints and the first displayed content, such as the target text displayed on the first interface, and is denoted as the third text content. It can be understood that the literary genre of this third text content is a story. In other embodiments, the text editing interface 220 may also include controls with other large-scale modeling capabilities, such as novel creation controls, fairy tale creation controls, etc., without limitation.
[0168] S305: The user has selected the first editing function in the first category of functions.
[0169] For example, mobile phone 100 can detect the user's operation of selecting a first editing function in the function list interface of the first type of functions. When the first editing function is displayed together with the aforementioned first type of functions in the aforementioned at least one editing function, mobile phone 100 can also detect the user's operation of selecting the first editing function in the aforementioned at least one editing function. The first editing function may include any function in the first type of functions, for example, it may include any one of the editing functions included in functions such as AI generation tool 212, such as writing assistant 221, poetry and prose 222, AI drawing 223, and story creation 224.
[0170] S306: Call the first generation model related to the first editing function and display the first editing interface.
[0171] For example, mobile phone 100 can provide the editing functions required by the memo application, such as those in the AI generation tool 212 mentioned above, register the relevant capabilities of the large model, and pre-build each generation model to implement the corresponding function. Each editing function can be implemented by calling the corresponding generation model. For example, the first editing function can be implemented by calling the first generation model. Taking the first editing function as writing assistant 221 as an example, the first generation model called by this first editing function can be a writing-type generation model pre-built based on the text expansion function of the large model. Similarly, for the first editing function as poetry 222, the first generation model called can be a retrieval-type generation model pre-built based on the database retrieval capability of the large model. For the first editing function as AI drawing 223 or story creation 224, the first generation model called can be a creation-type generation model pre-built based on the reasoning creation capability of the large model, and so on, without further elaboration. Based on this, mobile phone 100 can call the first generation model to provide the corresponding large model capabilities to implement the corresponding function when it detects that the user has selected the first editing function. Furthermore, based on the user interface (UI) capabilities of the first generative model, the mobile phone 100 can also display a first editing interface that allows the user to perform input operations.
[0172] The structure, function, and construction process of the generation models related to the various editing functions can be found in the following descriptions with accompanying figures, and will not be elaborated upon here.
[0173] As an example, corresponding to the editing function of writing assistant 221, the first editing interface described above can refer to the writing assistant interface 230 shown in Figure 2c. The writing assistant interface 230 can display parameter controls for controlling the final generated expanded content, such as "word count", "tone", and "stage" parameter controls.
[0174] S307: Detected that the user has set control parameters for at least one parameter adjustment control in the first editing interface.
[0175] For example, the mobile phone 100 can detect the user's operation of setting relevant parameters for one or more parameter adjustment controls on the first editing interface. Referring to Figure 2c above, this operation may include, for example, an operation corresponding to setting "word count" to "5000", an operation corresponding to setting "tone" to "professional", and an operation corresponding to setting "stage" to "university", etc.
[0176] It is understood that in other embodiments, corresponding to different first editing functions or editing functions provided by different target applications, the parameter adjustment controls displayed on the first editing interface may be different from the controls shown in Figure 2c above, and may also include more parameter adjustment controls than shown in Figure 2c, which is not limited here.
[0177] S308: Display at least one prompt related to setting control parameters in the first editing interface.
[0178] For example, based on the user's operation of setting relevant parameters for one or more parameter adjustment controls on the interface detected in step S307 above, the mobile phone 100 can generate corresponding prompts for the parameter adjustment results, such as "Word count: 5000", "Tone: Formal", "Stage: University", etc., and can display them in the description information input box of the first editing interface. The description information input box can refer to the style shown in the description information input box 232 in Figure 2c above. In other embodiments, the description information input box can also be displayed in other styles, which are not limited here.
[0179] S309: The user has confirmed the start of the generation operation on the first editing interface.
[0180] For example, after setting the various parameter controls on the first editing interface, the user can continue to instruct the generation of corresponding expanded content. For instance, referring to the writing assistant interface 230 shown in Figure 2c, the user can click the "Start Generation" control 231. Correspondingly, the mobile phone 100 can detect the user's click on the "Start Generation" control 231, that is, detect the user's instruction to generate corresponding expanded content on the first editing interface.
[0181] S310: Generate and display the first processing result based on the target text and at least one prompt word.
[0182] For example, the mobile phone 100 can convert the prompt words in the description information input box 232 into prompt information, and combine it with the target text selected by the user as input to the writing model to finally generate the expanded text content required by the user. At this time, the mobile phone 100 can display the expanded text content generated based on the target text and at least one prompt word, i.e., the first processing result.
[0183] It is understood that in other embodiments, the first processing result generated by the mobile phone 100 may also include the processed image, GIF, or video, etc., corresponding to the target image selected by the user; the first processing result generated by the mobile phone 100 may also include the edited audio, etc., corresponding to the target audio selected by the user; and the first processing result generated by the mobile phone 100 may also include the edited video, etc., without limitation.
[0184] S311: A user instruction to apply the first processing result to the first interface was detected.
[0185] For example, the interface on the mobile phone 100 displaying the first processing result may include one or more controls for further use or processing of the first processing result. Referring to the result generation interface 240 shown in FIG2d, the result generation interface 240 may include an application control 241, a copy control 242, a regenerate control 243, a share control 244, etc. Among them, the application control 241 can trigger the application of the currently displayed first processing result, such as expanded content 240a, to the first interface currently calling the text expansion function, such as the memo interface 210 shown in FIG2a. Correspondingly, if the mobile phone 100 detects a user's click operation on the application control 241 in the result generation interface 240, it detects an operation by the user instructing the application of the first processing result to the memo interface.
[0186] In other embodiments, the mobile phone 100 can also detect user clicks on the copy control 242, regenerate control 243, share control 244, etc., in the result generation interface 240, and can correspondingly execute the process of copying, regenerating, or sharing the first processing result described above. Further details are omitted here.
[0187] S312: Display a second interface including the results of the first processing.
[0188] For example, in response to a user instruction, the mobile phone 100 applies the first processing result to the memo interface. After applying the first processing result to the memo interface, a memo interface including the first processing result can be displayed, such as the memo interface 250 shown in Figure 2e above, which is the second interface of the memo application. Compared to the first interface of the memo application, such as the memo interface 210 shown in Figure 2a above, the first processing result displayed on the second interface can replace the target text selected by the user on the first interface, making the content of the corresponding note more complete.
[0189] In this way, users can apply the expanded text generated based on the large model capability to the target application that calls the large model capability without having to perform tedious operations such as copying, switching application interfaces, and pasting. This makes it convenient for users and helps improve the user experience.
[0190] The following section describes the specific implementation process of another data processing method provided in this application, using the application scenario of integrating the memo application call generation model with associated data to generate expanded content as provided in Example 2.
[0191] Example 2
[0192] Figure 4a illustrates an application scenario according to an embodiment of this application, where a memo application call generation model integrates associated data to generate expanded content. In this application scenario, the target application remains a memo application.
[0193] As shown in Figure 4a, the memo interface 410 displayed on the mobile phone 100 can show the note content that the user is editing, including a title 411 such as "My Family" and a paragraph of text 412 that has already been written, such as "There are five people in my family: Dad, Mom, younger brother, grandma, and me." The user needs to expand on the note content, for example, by writing an essay of about 800 words for elementary school students. Based on the data processing method provided in this application, the memo application can continue to call the above-mentioned generation model with text expansion function to provide functions such as "writing assistant" for users.
[0194] Unlike the application scenario applicable to Embodiment 1 above, the memo application in this application embodiment provides functions such as "writing assistant" and can integrate some related data generated by the user during the use of the mobile phone 100, such as related photos, preference settings and other data.
[0195] Based on the application scenario shown in Figure 4a above, Figure 5 illustrates a schematic diagram of the implementation process of a data processing method according to an embodiment of this application.
[0196] It is understood that, in this embodiment of the application, the executing entity for each step in the implementation process can continue to be the mobile phone 100. In other embodiments, the executing entity for each step in the implementation process shown in FIG5 can also be other electronic devices, such as tablet computers, laptop computers, etc., and there is no limitation here.
[0197] Specifically, the implementation process may include the following steps:
[0198] S501: In response to the user opening the Notes app, display the third interface of the Notes app.
[0199] For example, the third interface of the memo application may include the memo interface 410 shown in Figure 4a above. The memo interface 410 may, for example, display a composition note titled "My Family" that the user is editing.
[0200] S502: Detecting that the user has selected the target text on the third interface, display at least one editing function on the third interface.
[0201] For example, the user's operation of selecting target text on the third interface may include, for instance, the user long-pressing text in the current note content on the memo interface to select part or all of the text in the current note content as the target text. Referring to Figure 4a above, if the mobile phone 100 detects that the user long-presses text 412 on the memo interface 410, it can display at least one editing function on the memo interface 410, such as the various editing functions in the function option box 413, including AI generation tool 414, writing assistant 415, poetry and prose 416, etc. The function option box 413 can also display some editing functions that are not based on AI models, such as the "copy" function 417, the "cut" function 418, etc., for the user to select and use according to actual needs.
[0202] S503: The user has selected at least one of the first type of editing functions.
[0203] For example, the first type of function in at least one of the above-mentioned editing functions may include, for example, the AI generation tool 414. It is understood that the first type of function may include one or more editing functions that process data corresponding to a certain modality. Referring to Figure 4b, the multiple editing functions included in the AI generation tool 414 may include writing assistant 221, poetry and prose 222, AI drawing 223, and story creation 224, etc., and are not limited here. For a detailed description of each editing function, please refer to the relevant description of the interface shown in Figure 2b above, which will not be repeated here.
[0204] S504: Displays the function list interface for the first type of function.
[0205] For example, in response to the user's selection of at least one editing function from the first category, the mobile phone 100 can display a function list interface including all editing functions in the first category, such as the text editing interface 420 shown in Figure 4b. This text editing interface 420 can display writing assistant 421, poetry and prose 422, AI drawing 423, and story creation 424, etc. For a detailed description of the function list interface, please refer to the description of the interface shown in Figure 2b above; it will not be repeated here.
[0206] S505: The user has selected the first editing function in the first category of functions.
[0207] The specific execution of step S505 can be found in the description of step S305 in the above embodiment 1, and will not be repeated here.
[0208] In this embodiment of the application, the first editing function in the first type of function can still be any editing function in the text editing interface 420 shown in FIG4b, such as the writing assistant 421.
[0209] S506: Call the first generation model related to the first editing function and display the second editing interface.
[0210] For example, taking the writing assistant 421 as an example, the mobile phone 100 can respond to the user's selection of the writing assistant 421, call the pre-built generation model based on the large model text expansion function, and display the writing assistant interface 430, i.e., the second editing interface, as shown in Figure 4c.
[0211] Unlike the first editing interface displayed in step S306 of Embodiment 1 above, the writing assistant interface 430 shown in FIG4c can include not only a "Start Generation" control 431, a description information input box 432, and parameter adjustment controls such as "Word Count" and "Stage", but also an association control 433. In this embodiment, the association control 433 can be local data from electronic devices such as mobile phones. This local data can include application data from other applications, such as image or video data from a gallery application, calendar data from a calendar application, map data from a map application, email data from an email application, text data from a notepad or text editing application, audio data from a recording application, contact data, and SMS data. Taking image or video data from a gallery application as an example, the association control 433 indicates association with gallery people, corresponding to the "Gallery People Association" shown in FIG4c.
[0212] Referring again to Figure 4c, after the user turns on the associated control 433, the writing assistant interface 430 can display related character categories 434 associated with "My Family". These categories 434 can be distinguished using user-defined roles or identities, such as "Dad", "Mom", or "Brother" set from the perspective of one of the children. Correspondingly, the related character category 434 can display data including those labeled "Dad", "Mom", and "Brother", such as a photo album 434a containing photos or videos of the father, or a photo album 434b containing photos or videos of the mother. In some embodiments, the user can also set roles or identities from their own perspective, such as "I", "Wife", "Eldest Child", or "Second Child", which will not be elaborated upon here. In some embodiments, the related character category 434 can also include a category of group photos reflecting the relationships between the characters, such as the family photo album 434c shown in Figure 4c, which is not limited here. In some embodiments, if there are no personal photos of the associated person in the local image or video data of the electronic device such as mobile phone 100, the group photo in the above-mentioned image set 434c can provide relevant features of the associated person to generate relevant feature prompts for the person, i.e., prompt information.
[0213] It is understandable that before calling the first generative model, the mobile phone 100 can perform semantic understanding on the aforementioned associated data to obtain semantic understanding results. Based on these results, keywords describing the associated data are extracted as prompt information and input into the subsequently called first generative model as control parameters for generating the corresponding processing results. Given that the associated data may include modalities such as text, images, or videos, the semantic understanding results may include semantic understanding results for text, or image recognition results for images or videos. These image recognition results may extract image feature data from people, scenery, etc., in the relevant images or videos. The mobile phone 100 can further convert this image feature data into prompt information describing the image features of the relevant people, scenery, etc., and input this information into the subsequently called first generative model. Ultimately, the processing results output by the first generative model can include information obtained from the aforementioned associated data, such as person feature information, person relationship information, and environmental feature information.
[0214] It is understood that the associated control 433 shown in Figure 4c above takes the form of a switch as an example. Users can control whether the "Gallery People" indicated by the associated control 433 is associated or not by clicking the switch to either the on or off state. In other embodiments, the associated control may also take other forms besides a switch, which is not limited here. Examples of the associated data being other application data (such as calendar data, map data, etc.) will be exemplified below with relevant interface diagrams, and will not be elaborated upon here. In other embodiments, the associated control 433 may also be associated with user data related to user collection or liking behaviors in apps such as Xiaohongshu™, which is not limited here.
[0215] S507: Detected user operation in the second editing interface to set control parameters for at least one parameter adjustment control.
[0216] For example, continuing to refer to FIG4c, the difference from the execution of step S307 in Embodiment 1 above is that the user's operation of setting control parameters for at least one parameter adjustment control in the second editing interface can include not only setting the parameter adjustment control, such as "word count" to "800" and "stage" to "primary school", but also the operation of selecting associated persons in associated person photos 434.
[0217] In this embodiment, the associated controls corresponding to the aforementioned associated person photos 434 can enable target applications, such as memo applications, to obtain associated user behavior data, as well as relevant feature vectors or feature clustering results pre-extracted from the user behavior data. The specific process by which the associated controls implement their corresponding associated functions can be referred to in the description of Figure 12d and related content below, and will not be elaborated upon here.
[0218] In some embodiments of this application, electronic devices such as mobile phones 100 can also display associated data input boxes or input controls for users to input information through a writing assistant interface 430 as shown in Figure 4c. The input information received by the user corresponding to the input box or input control can be processed by keyword extraction and other methods to search for associated data from associated local data already stored in the electronic device such as mobile phones 100 or feature vectors extracted from local data. Then, the searched associated data is combined with the user's operation results of setting control parameters for other parameter adjustment controls, and used together to execute the following steps S508 to S509 to generate prompt information and generate the second processing result required by the user.
[0219] S508: Display at least one prompt related to setting control parameters in the second editing interface.
[0220] For example, continuing to refer to Figure 4c, the difference from the execution of step S308 in the above embodiment 1 is that, based on the operation of setting control parameters for at least one parameter adjustment control detected in step S507, the mobile phone 100 can display prompts corresponding to each parameter adjustment result in the description information input box 432 of the writing assistant interface 430, such as "word count: 800", "stage: primary school", etc. In some embodiments, it can also display words such as "related person.jpg", which is not limited here.
[0221] S509: The user has confirmed the start of the generation operation on the second editing interface.
[0222] The specific execution of step S509 can be found in the description of step S309 in the above embodiment 1, and will not be repeated here.
[0223] It is understandable that the mobile phone 100 responds to the user's confirmation of starting the generation operation on the second editing interface, such as the writing assistant interface 430, and may display the generation result interface 440 as shown in Figure 4d.
[0224] S510: Generate and display the second processing result based on the target text and at least one prompt word.
[0225] For example, the second processing result mentioned above can refer to the expanded content 440a displayed in the result generation interface 440 shown in Figure 4d, such as: "There are five people in my family: my dad, mom, younger brother, grandma, and me. We love each other very much, and our home is full of laughter. My dad is tall and thin with thick, short hair. He likes sports and often takes us camping, cycling, and playing soccer together, which is very fun. There is someone who is like the sun in our family, always taking care of me, this little sapling; that is my mom. My mom has fair skin, long hair, and a pair of dark eyes. She loves to laugh. She is always smiling, which makes people feel very kind and warm. We also have a 3-year-old troublemaker in our family. He is very naughty. He often paints his face in all sorts of colors and makes some very 'creative' shapes."
[0226] In this embodiment of the application, the second processing result, such as the expanded content 440a, can incorporate image description information corresponding to the associated person photos provided by the associated person photo 434, such as "Dad is tall and thin," often taking us to "ride bicycles and play soccer," "Mom has fair skin and long hair," "My younger brother paints his face in various colors," and likes to make "creative" shapes, etc. This expanded content 440a, i.e., the second processing result, can better meet user expectations; for example, the expanded content 440a can better meet the user's writing requirements.
[0227] It is understandable that the mobile phone 100 can convert the prompt words in the above description information input box 432 into prompt information, and can also convert the image description information corresponding to the above-mentioned associated person photos, such as the physical characteristics of "father" being "tall" and "thin", and the physical characteristics of "mother" being "fair" and "long hair", into prompt information, and combine them with the target text selected by the user as input to the writing model, and finally generate the text expansion content required by the user.
[0228] S511: A user instruction to apply the second processing result to the third interface was detected.
[0229] For example, the interface on the mobile phone 100 displaying the second processing result may include one or more controls for further use or processing of the second processing result. Referring to the result generation interface 440 shown in FIG4d, the result generation interface 440 may include an application control 441, a copy control 442, a regenerate control 443, a share control 444, etc. If the mobile phone 100 detects a user's click operation on the application control 441 in the result generation interface 440, it detects an operation by the user instructing the second processing result to be applied to a third interface (e.g., the memo interface 430 mentioned above).
[0230] In other embodiments, the mobile phone 100 can also detect user clicks on the copy control 442, regenerate control 443, share control 444, etc., in the result generation interface 440, and can correspondingly execute the process of copying, regenerating, or sharing the second processing result described above. For details, please refer to the relevant description in Figure 2d of step S311 in Embodiment 1 above, which will not be repeated here.
[0231] S512: Display a fourth interface including the second processing result.
[0232] For example, in response to a user's instruction to apply the second processing result to the memo interface, the mobile phone 100 can display a memo interface including the second processing result, such as the memo interface 450 shown in Figure 4e, which is the fourth interface of the memo application. Unlike the second interface displaying the first processing result in step S312 of Embodiment 1, the second processing result displayed by the mobile phone 100, such as the memo interface 450 shown in Figure 4e, can include expanded content 440a generated from the notes corresponding to the title "My Family." As mentioned earlier, this expanded content 440a can incorporate image description information corresponding to the associated person photos provided by the associated person photos 434, such as "Dad is tall and thin," often taking us to "ride bicycles and play soccer," "Mom has fair skin and long hair," "My younger brother paints his face in various colors," and likes to make "creative" shapes, etc. Therefore, this expanded content 440a better meets the user's writing needs or text expansion needs for an essay on the theme "My Family."
[0233] For an example where the associated data involved in step S506 above is other application data (such as schedule data, map data, etc.), the following is an exemplary description in conjunction with Figures 4f and 4g.
[0234] Referring to the schedule interface 460 shown in Figure 4f, the schedule data that can be used as associated data can include information such as time and events, and can also include plans based on one or more events, etc., without limitation. Time information includes, for example, the calendar information 461a and event-related time information 461b shown in Figure 4f. Time information 461b can include a specific time period of a specific day in calendar information 461a, such as the event time period "11:00-11:10" reminded by schedule card 462 shown in Figure 4f. Event information can include the events reminded by schedule card 462 shown in Figure 4f. Event content can include location and event theme, such as an event or concert held at the "Nanjing Olympic Sports Center Stadium".
[0235] Referring to the map interface 470 shown in Figure 4g, the map data that can be used as associated data may include map check-in data based on location records, such as "Light up the city" 471, "Check-in favorites" 472, "Travel mileage" 473, etc. This map check-in data can record the locations the user has visited, the time of arrival at the corresponding location, and the time period of stay, etc., without any restrictions.
[0236] Other application data are examples of related data, and will not be described further in this application embodiment.
[0237] The following section continues with the application scenario provided in Example 3, which uses a memo application to call a generative model with image processing capabilities to generate a virtual image, to introduce the specific implementation process of another data processing method provided in this application.
[0238] Example 3
[0239] Figure 6a illustrates an application scenario according to an embodiment of this application, where a memo application calls a generative model with image processing capabilities to generate a virtual avatar. In this application scenario, the target application remains a memo application. For example, the memo application may need to change the style of images in related notes, processing them into virtual avatars with a specified style.
[0240] As shown in Figure 6a, the memo interface 610 displayed on the mobile phone 100 may include an image 611. The user needs to process this image into a specific virtual avatar corresponding to the person in the image. For example, the processed virtual avatar may be dressed in clothing of a specific style.
[0241] Based on the application scenario shown in Figure 6a, Figure 7 illustrates a schematic diagram of the implementation process of a data processing method according to an embodiment of this application.
[0242] It is understood that, in this embodiment of the application, the executing entity for each step in the implementation process can continue to be the mobile phone 100. In other embodiments, the executing entity for each step in the implementation process shown in FIG7 can also be other electronic devices, such as tablet computers, laptop computers, etc., and there is no limitation here.
[0243] Specifically, the implementation process may include the following steps:
[0244] S701: In response to the user opening the Notes app, display the fifth screen of the Notes app.
[0245] The specific execution process of the mobile phone 100 responding to user operations and opening the fifth interface of the memo application can be referred to the relevant description in step S301 of the above embodiment 1, and will not be repeated here.
[0246] Unlike the execution of step S301 above, the memo interface opened by the mobile phone 100 in step S701 may include an image interface, such as the memo interface 610 shown in Figure 6a above. For ease of distinction, interfaces such as memo interface 610 may be referred to below as the fifth interface of the memo application.
[0247] S702: Detects that the user has selected a target image on the fifth interface, and displays at least one editing function on the fifth interface.
[0248] For example, the content of the current note displayed on the memo interface (i.e., the fifth interface mentioned above) may include an image, and the content selected by the user is the target image. The mobile phone 100 can detect when the user long-presses an image in the current note content on the memo interface, selecting one or more images from the current note content as the target image, and then display one or more editing functions that allow editing of the selected target image.
[0249] It is understood that at least one editing function displayed on the fifth interface can be various editing functions corresponding to the target image. These editing functions can include a second category of functions related to image editing or image processing. These second category of functions can be implemented by calling corresponding generative models built based on large model capabilities, i.e., based on AI models.
[0250] As an example, Figure 6b illustrates a schematic diagram of a memo interface including at least one editing function according to an embodiment of this application.
[0251] As shown in Figure 6b, the memo interface 620 can display a function option box 621 corresponding to the image selected by the user. This function option box 621 can include a second type of function, such as an AI generation tool 622, and some editing functions included in this second type of function, such as image editing 623 and poster generation 624 in the function option box 621 shown in Figure 6b.
[0252] In addition, at least one of the above editing functions may also include some editing functions that are not based on AI models, such as the "copy" function 625 and the "cut" function 626 in the function option box 621 shown in Figure 6b, which users can choose to use according to their actual needs.
[0253] S703: The user has been detected to have selected at least one of the second type of editing functions.
[0254] For example, the second type of function selected by the user in at least one editing function displayed on the mobile phone 100 may include, for example, the AI generation tool 622 in the function option box 621 shown in Figure 6b above. The mobile phone 100 detects the user's click on the AI generation tool 622, that is, detects the user's selection of the second type of function in at least one editing function. The second type of function may include one or more editing functions, each of which can be implemented by the memo application calling the corresponding generation model. For details, please refer to the relevant description below, which will not be repeated here.
[0255] S704: Displays the function list interface for the second category of functions.
[0256] For example, in response to the detected user selection of the second type of function, such as the user clicking the AI generation tool 622 in the memo interface 620 shown in Figure 6b above, the mobile phone 100 can display the various editing functions included in the second type of function. This display interface is described in this application embodiment as a function list interface of the second type of function.
[0257] As an example, the various editing functions included in the AI generation tool 622 can be seen in Figure 6c. In response to the user clicking the AI generation tool 622 on the memo interface 620, the mobile phone 100 can display the image editing interface 630. As shown in Figure 6c, the image editing interface 630 may include image editing 631, poster generation 632, high-definition restoration 633, local modification 634, image expansion 635, style change 636, and background blurring 637, among other function options.
[0258] The image editing component 631 can be a control used to trigger image processing functions. These image processing functions are used to edit the foreground subject of the target image according to the aforementioned constraints and the first displayed content, such as the target image displayed on the fifth interface of the memo application. For example, they can perform beautification processes on the subject, such as skin smoothing and retouching. The final processed image can be referred to as the second image. It can be understood that the foreground subject in the first image and the second image is the same.
[0259] Poster generator 632 can be a control used to trigger the function of generating a poster. This poster generation function processes the target image into a poster with promotional effects based on the aforementioned constraints and the first displayed content, such as the target image displayed on the fifth interface of the memo application. The final processed image can be referred to as the third image.
[0260] The HD Repair 633 is a control used to trigger a function to repair the clarity of images or videos. This function can use image noise reduction algorithms or other image processing algorithms to perform image processing procedures, including noise reduction, on the target image displayed in the fifth interface of the aforementioned memo application, ultimately resulting in a clearer image or video.
[0261] The partial modification 634 can be a control used to trigger partial modification, partial redraw, or region modification capabilities. This partial modification, partial redraw, or region modification capability can modify the color or redraw details of a portion of the target image displayed on the fifth interface of the aforementioned memo application, such as the first displayed content.
[0262] Image expansion 635 can be a control used to trigger the image expansion function. This image expansion function, based on the aforementioned constraints, expands the size of the background area of the first displayed content, such as the first image displayed on the fifth interface of the aforementioned memo application, to obtain a distant image where the foreground subject appears visually farther from the background.
[0263] Style change 636 can be a control used to trigger a function to change the style of an image or video. This function, which changes the style of an image or video, can change the style of the first displayed content, such as the foreground subject or background of the target image displayed on the fifth interface of the memo application, according to the constraints mentioned above, such as the "future" style mentioned below.
[0264] Background blur 637 can be a control used to trigger the background blur processing function. This background blur processing function is used to blur the background of the first displayed content, such as the target image displayed on the fifth interface of the memo application, according to the above constraints, in order to highlight the foreground subject.
[0265] Referring again to Figure 6c, the image editing interface 630 may also include a description information input box 638, which may display prompts such as "Please describe your needs". Users can input their processing requirements for the image 611 in the aforementioned memo interface 610 into this description information input box 638.
[0266] In this embodiment, the user can also select image editing 631 as the basic editing function and choose to overlay other editing functions as specific image editing requirements, inputting them to the corresponding generative model with image processing capabilities. In this case, the user can set control parameters in the options corresponding to the other editing functions, which can also be displayed in the description information input box 638 as processing requirements for the corresponding image. As mentioned earlier, the generative model with image processing capabilities can include UI components built based on the capabilities of the corresponding large model. The other editing functions selected by the user can be correspondingly generated as prompt information input to the corresponding large model to control the processing results, which will not be elaborated upon here.
[0267] S705: The user has selected the second editing function in the second category of functions.
[0268] For example, the second editing function in the second category of functions mentioned above may include one or more of the functions of image editing 631, poster generation 632, high-definition restoration 633, local modification 634, image expansion 635, style change 636, and background blurring 637 belonging to the AI generation tool 622, and there is no limitation here. Referring to Figure 6c above, for example, a user can select image editing 631 in the image editing interface 630, or a user can first select image editing 631 and then select style change 636 in the image editing interface 630 as an overlay editing function, etc. Correspondingly, the mobile phone 100 can detect the user's operation of selecting the second editing function in the second category of functions.
[0269] S706: Invoke the second generation model related to the second editing function and display the third editing interface.
[0270] For example, in response to a user selecting a second editing function, the mobile phone 100 can invoke a second generation model corresponding to the second editing function to implement the second editing function. Furthermore, based on the UI capabilities of the second generation model, the mobile phone 100 can display a third editing interface corresponding to the second editing function.
[0271] As an example, the third editing interface described above can refer to the image editing interface 640 shown in Figure 6d. As shown in Figure 6d, taking the user's selection of overlaying a higher style 636 on image editing 631 as an example, the image editing interface 640 may include a description information input box 641 and multiple styles 642 available for the user to choose from.
[0272] S707: Detected user operation in the third editing interface to set control parameters for at least one parameter adjustment control.
[0273] For example, the operation of setting control parameters for at least one parameter adjustment control in the third editing interface may include the operation of the user selecting the overlay editing function in the image editing interface 630 shown in Figure 6c above, or the operation of the user selecting a preferred style in the corresponding overlay style change 636 in the image editing interface 640, etc., without limitation.
[0274] In other embodiments, the operation of setting control parameters for at least one parameter adjustment control in the third editing interface may also include the operation of the user selecting or setting relevant parameters for editing functions such as poster generation 632, high-definition restoration 633, local modification 634, image expansion 635, and background blurring 637, which is not limited here.
[0275] S708: Display at least one prompt related to setting control parameters in the third editing interface.
[0276] For example, continuing to refer to Figure 6d, if the user selects the overlay style change 636 and selects the "future" style, the mobile phone 100 can correspondingly display prompts such as "Style Change: Future" in the description information input box 641 of the image editing interface 640.
[0277] S709: The user has confirmed the start of the generation operation on the third editing interface.
[0278] For example, after the user completes the setting operations for various parameter control on the third editing interface, they can continue to instruct the user to generate the desired image. For instance, referring to the image editing interface 640 shown in Figure 6d, the user can click the "Start Generating" control 643. Correspondingly, the mobile phone 100 can detect this click on the "Start Generating" control 643, that is, detect the user's instruction to generate the desired image on the third editing interface.
[0279] S710: Generate and display a third processing result based on the target image and at least one prompt word.
[0280] For example, the mobile phone 100 can convert the prompt words in the description information input box 641 into prompt information, and combine it with the target image selected by the user by long-pressing in the memo interface 610 as input to the generative model with image processing capabilities, and finally generate the image required by the user. At this time, the mobile phone 100 can display the processed image generated based on the target image and at least one prompt word, which is recorded as the third processing result.
[0281] As an example, the interface displaying the third processing result can be seen in Figure 6e. That is, after the mobile phone 100 detects the user's operation of clicking the "Start Generation" control 643 on the image editing interface 640, it can continue to display the generation result interface 650 shown in Figure 6e.
[0282] As shown in Figure 6e, the generated result interface 650 can display an image 650a generated based on the target image and at least one prompt word, such as a virtual avatar image in a "futuristic" style. To facilitate user style changes, the generated result interface 650 also allows users to select other alternative styles.
[0283] S711: A user instruction to apply the third processing result to the fifth interface has been detected.
[0284] For example, continuing to refer to Figure 6e above, the result generation interface 650 may further include an application control 651, a copy control 652, a regenerate control 653, a share control 654, etc. The application control 651 can trigger the application of the currently displayed third processing result, such as expanded content 650a, to the memo application currently calling the image processing function. Correspondingly, if the mobile phone 100 detects a user's click operation on the application control 651 in the result generation interface 650, it detects an operation instructing the user to apply the third processing result to the memo interface.
[0285] In other embodiments, the mobile phone 100 can also detect user clicks on the copy control 652, regenerate control 653, share control 654, etc., in the result generation interface 650, and can correspondingly execute the process of copying, regenerating, or sharing the third processing result described above. Further details are omitted here.
[0286] S712: Displays the sixth interface, which includes the third processing result.
[0287] For example, in response to a user instruction, the mobile phone 100 applies the third processing result to the memo interface. After applying the third processing result to the memo interface, a memo interface including the third processing result can be displayed, such as the memo interface 660 shown in Figure 6f above, which is the second interface of the memo application. Compared to the second interface of the memo application, such as the memo interface 610 shown in Figure 6a above, the third processing result displayed on the second interface can replace the target image selected by the user on the second interface, making the content of the corresponding note more complete.
[0288] In this way, users can apply the images generated based on the large model capabilities to the target application that calls the large model capabilities without having to perform cumbersome operations such as saving images, adding saved images to the relevant interface, saving / exporting generated images, etc. This makes it convenient for users and helps improve the user experience.
[0289] The following section continues with the application scenario provided in Example 4, where a generative model with image processing capabilities is used to change the clothing of people in an image to generate a virtual avatar. This section introduces the specific implementation process of another data processing method provided in this application.
[0290] Example 4
[0291] Figure 8a illustrates an application scenario according to an embodiment of this application, where a gallery application calls a generative model with image processing capabilities to generate a virtual avatar. In this application scenario, the target application is a gallery application.
[0292] As shown in Figure 8a, the gallery interface 810 displayed on the mobile phone 100 can show the photos of people currently viewed by the user. In some scenarios, the user may want to quickly change the clothing of the person in the photo, such as adding a hat, wearing a watch or headphones, or changing the outfit. In this case, based on the data processing method provided in this application, the mobile phone 100 can call a generative model with corresponding image processing functions when running the gallery application to provide the user with image processing functions for quickly changing the person's clothing.
[0293] Based on the application scenario shown in Figure 8a above, Figure 9 illustrates a schematic diagram of the implementation process of a data processing method according to an embodiment of this application.
[0294] It is understood that, in this embodiment of the application, the executing entity for each step in the implementation process can continue to be the mobile phone 100. In other embodiments, the executing entity for each step in the implementation process shown in FIG9 can also be other electronic devices, such as tablet computers, laptop computers, etc., and there is no limitation here.
[0295] Specifically, the implementation process may include the following steps:
[0296] S901: In response to the user opening the Gallery app, display the seventh screen of the Gallery app.
[0297] For example, the seventh interface of the aforementioned gallery application can be the gallery interface 810 shown in FIG8a. Referring to FIG8a, the images displayed in the gallery interface 810 can be, for example, the photos of people currently viewed by the user.
[0298] S902: The user's action of selecting the target image on the seventh interface is detected, and at least one editing function is displayed on the seventh interface.
[0299] For example, the user's operation of selecting a target image on the seventh interface may include, for instance, the user long-pressing a person's photo displayed on the gallery interface 810 shown in Figure 8a. Referring to Figure 8a, the user can long-press the person's photo on the gallery interface 810. Correspondingly, in response to this long-press operation, the mobile phone 100 can display at least one editing function, that is, display at least one editing function on the seventh interface.
[0300] Referring again to Figure 8a, in the function option box 811 displayed on the gallery interface 810, at least one editing function may include AI generation tools 812, clothing and styling 813, and style changing 814, etc. These editing functions can be implemented based on AI models. For example, each editing function can call a generation model built based on the capabilities of a large model to implement the corresponding editing function. Among them, clothing and styling 813 can be a control that has the function of changing the clothing and styling of people in pictures or videos.
[0301] In other embodiments, the mobile phone 100 may also display the aforementioned AI generation tools and other editing functions in the more controls 815 on the gallery interface 810 for users to choose from, without any limitation.
[0302] S903: The user has been detected to have selected at least one of the second type of editing functions.
[0303] For example, the second type of function may include the various image processing functions provided by the corresponding generation model called by the memo application in Embodiment 3 above, such as the various editing functions included in the AI generation tool 812 shown in Figure 8a. In this application embodiment, the above-mentioned AI generation tool and other second type of functions may differ from the image processing functions already available in the gallery application, such as supporting cropping, adjustment (brightness, contrast, saturation, etc.), filters, and doodles. In this application embodiment, the above-mentioned second type of function refers to the function of extracting human features from images, modifying clothing and outfits, adjusting human posture, and switching styles based on AI models to create secondary images.
[0304] As an example, referring to Figure 8a, the editing functions included in the AI generation tool 812 may include clothing matching 813 and style changing 814. In response to the user clicking the AI generation tool 812 and selecting the clothing matching option, or clicking the clothing matching 813 option, on the gallery interface 810 shown in Figure 8a, the mobile phone 100 can display the AI generation interface 820 shown in Figure 8b.
[0305] S904: Displays the function list interface for the second category of functions.
[0306] For example, the function list interface of the second type of function in this application embodiment may include, for example, the various editing functions selectable in the function option box 821 of the AI generation interface 820 shown in FIG8b, including the above-mentioned functions such as clothing matching and style change.
[0307] S905: The user was detected selecting the third editing function in the second category of functions.
[0308] For example, the third editing function in this application embodiment may include any one of the functions of clothing matching and style change belonging to the AI generation tool 812 (i.e., the second type of function). As an example, the operation of the user selecting the third editing function in the second type of function may be the operation of clothing matching 813 displayed on the gallery interface 810 shown in Figure 8a, or the operation of setting the function in the function option box 821 of the AI generation interface 820 shown in Figure 8b to the clothing matching function, etc., without limitation.
[0309] S906: Invoke the third generation model related to the third editing function and display the fourth editing interface.
[0310] For example, the third generative model associated with the third editing function can be a generative model pre-built by the mobile phone 100 using the image processing function of a large model for editing and creating target images. Based on the third generative model associated with the third editing function, the mobile phone 100 can also display a fourth editing interface including at least one parameter adjustment control. In some embodiments, when the mobile phone 100 calls the third generative model, it can also load some pre-trained clothing matching models and provide parameter adjustment controls on the displayed fourth editing interface for users to set relevant parameters, which is not limited here.
[0311] As an example, the fourth editing interface described above can be referenced to the AI generation interface 820 shown in Figure 8b. At least one parameter adjustment control displayed on this fourth editing interface can be referenced from the parameter adjustment controls for personal size, style options, and clothing material in the AI generation interface 820 shown in Figure 8b. For a detailed description of the parameter adjustment controls on this fourth editing interface, please refer to the relevant description of Figure 8b below; it will not be repeated here.
[0312] S907: Detected user operation in the fourth editing interface to set control parameters for at least one parameter adjustment control.
[0313] For example, referring to the AI generation interface 820 displayed on the mobile phone 100 shown in Figure 8b, it may include multiple parameter adjustment controls in the parameter adjustment control area 823 related to parameters such as "personal size," "style options," and "clothing material." For instance, a user can use the parameter adjustment controls related to "personal size" in the parameter adjustment control area 823 to set parameters such as "height," "chest," "waist," and "hips." As another example, a user can use the parameter adjustment controls related to "style options" in the parameter adjustment control area 823 to set parameters such as "sports" to control the recommended clothing style. As yet another example, a user can use the parameter adjustment controls related to "clothing material" in the parameter adjustment control area 823 to set parameters such as "pure cotton" to control the recommended clothing.
[0314] In other embodiments, parameters such as “personal size,” “style options,” and “clothing material” can also be obtained through associated data. For example, when the gallery application running on mobile phone 100 provides the various editing functions of the AI generation tool, in response to the user’s authorization to obtain associated data from other applications, it can obtain relevant parameters such as the user’s personal size, style preferences, and material preferences set in some shopping applications, without limitation.
[0315] Referring again to Figure 8b, after the user sets control parameters for at least one parameter adjustment control in the fourth editing interface, the description information input box 822 in the AI generation interface 820 can display prompts corresponding to each set control parameter, including prompts for parameters such as personal size, style options, and clothing material. Furthermore, this description information input box 822 allows the user to input descriptive information about clothing combinations.
[0316] In some embodiments, the "View More Options" 824 in the AI generation interface 820 shown in Figure 8b can also allow users to set more other types of parameters related to clothing and outfits, which are not limited here.
[0317] S908: Display at least one prompt related to setting control parameters in the fourth editing interface.
[0318] For example, in the fourth editing interface, such as the description information input box 822 of the AI generation interface 820 shown in Figure 8b above, the description information of clothing and outfit entered by the user can be displayed, as well as prompts corresponding to different types of parameters set by the user in the parameter adjustment control area 823.
[0319] S909: The user has confirmed the start of the generation operation on the fourth editing interface.
[0320] For example, the user confirms the start of generation on the fourth editing interface, which may include the user clicking the "Start Generation" control 825 in the AI generation interface 820 shown in Figure 8b above.
[0321] S910: Generate and display the fourth processing result based on the target image and at least one prompt word.
[0322] For example, referring to FIG8c, based on the user operation detected in step S909, the mobile phone 100 displays the generated image that meets the user's needs on the result generation interface 830, such as a portrait photo 830a after changing the person's clothing. Continuing to refer to FIG8c, the result generation interface 830 can also display various types of clothing applied to the portrait photo 830a, such as hats, watches, glasses, headphones, etc. The result generation interface 830 can also display other clothing recommendations for the user to choose from, as well as application controls 831, copy controls 832, regenerate controls 833, share controls 834, etc.
[0323] S911: A user instruction to apply the fourth processing result to the seventh interface has been detected.
[0324] S912: Displays the eighth interface, which includes the fourth processing result.
[0325] For example, if the mobile phone 100 detects a user's click operation on the application control 831 in the generated result interface 830, it can apply the generated portrait photo 830a to the gallery application and display the gallery interface 840 shown in Figure 8d.
[0326] It is understandable that the description information and prompts displayed in the fourth editing interface, such as the description information input box 822 shown in Figure 8b, can all be converted into prompt information based on the prompt generation model integrated during the construction of the third generation model. This information, along with the user-specified edited photo, is then input into the large model or AI model upon which the third generation model relies for model inference. In this case, the prompt generation model of the third generation model, based on the prompt information generated by the user through at least one of the aforementioned parameter tuning controls, may be more accurate than the prompt information manually entered by the user in the description information input box. Furthermore, combining the user's input in the description information input box with the parameters set by the user through at least one of the aforementioned parameter tuning controls can generate even more accurate prompt information. Thus, more accurate prompt information allows for more precise matching of the generated results to user needs, improving the user experience.
[0327] Furthermore, the processing results obtained by the mobile phone 100 based on the data processing method provided in this application can be directly applied to the relevant interface of the target application that calls the corresponding function based on the AI model. This simplifies the process for users to perform complex operations such as copying, pasting, saving, or loading the processing results in the AI application and then apply them to the target application, which also helps to improve the user experience.
[0328] The following section continues with the application scenario provided in Example 5, which uses a generative model with video editing capabilities to generate storyboard images for advertising storyboards, and introduces the specific implementation process of another data processing method provided in this application.
[0329] Example 5
[0330] Figure 10a illustrates an application scenario according to an embodiment of this application, where a creative application calls a generative model with video editing capabilities to generate storyboard images for an advertisement. In this application scenario, the target application is a creative application, which may include video editing applications such as Huaban Clip™, CapCut™, or NewFilm™.
[0331] As shown in Figure 10a, the creation interface 010 displayed on the mobile phone 100 can include multiple creation functions. Each creation function can extract corresponding types of creative materials based on the target video specified by the user for creation. These creation functions include "Ad Storyboard" 011, as well as functions such as wallpaper inspiration, story inspiration, and proposal inspiration. For example, if a user needs to use the "Ad Storyboard" creation function to generate creative storyboard images, they can click on the "Ad Storyboard" 011 function option to create the storyboard.
[0332] Based on the application scenario shown in Figure 10a, Figure 11 illustrates a schematic flowchart of a data processing method according to an embodiment of this application. Based on the flowchart shown in Figure 11, the mobile phone 100 can call an AI generation tool to quickly complete the storyboard production process and provide the user with a relatively satisfactory storyboard production result.
[0333] It is understood that, in this embodiment of the application, the executing entity for each step in the implementation process can continue to be the mobile phone 100. In other embodiments, the executing entity for each step in the implementation process shown in FIG11 can also be other electronic devices, such as tablet computers, laptop computers, etc., and there is no limitation here.
[0334] Specifically, the implementation process may include the following steps:
[0335] S1101: In response to the user opening the creation application, the ninth interface of the creation application is displayed.
[0336] For example, mobile phone 100 can respond to a user clicking the icon of the creation application, run the creation application, and display the main interface or other related interfaces of the creation application. The creation application may include a generative app capable of creating advertising storyboards, wallpapers, stories, and creative proposals, or providing a platform for users to create various types of content. In some embodiments, the creation application may also be an image enhancement application, a video editing application, etc., without limitation.
[0337] S1102: Display at least one editing function on the ninth interface.
[0338] For example, referring to Figure 10a above, at least one editing function displayed on the ninth interface may be, for example, the various editing functions in the function option box 012 displayed in the creation interface 010 of the mobile phone 100. In response to the user clicking the "Advertisement Storyboard 011" function option in the creation interface 010, the mobile phone 100 may display the function option box 012 related to creating advertisement storyboards. The "AI Generation Tool" in this function option box 012 can provide advertisement storyboard creation functions based on AI models.
[0339] S1103: The user was detected to have selected the fourth editing function among at least one editing function.
[0340] For example, referring to Figure 10a above, when a user selects "AI Generation Tool" in the function option box 012 displayed in the creation interface 010, the user can select the fourth editing function among the above-mentioned at least one editing function. Correspondingly, this fourth editing function can be the function of generating storyboard images provided by the "AI Generation Tool" corresponding to the "Advertising Storyboard" function.
[0341] In other embodiments, the aforementioned "AI generation tool" can also be a type of function, for example, it can be referred to as a third type of function. Referring to Figure 10a, unlike the first and second types of functions mentioned above, the "AI generation tool" can include editing functions for images or videos corresponding to the currently selected ad scene 011, and can also include image or video editing functions corresponding to functions such as "wallpaper inspiration," "story inspiration," and "proposal inspiration" shown in Figure 10a, etc., without limitation.
[0342] Among them, the advertising storyboard 011 can be a control used to trigger the function of creating storyboard images. This function of creating storyboard images can perform image cropping, adding relevant elements for advertising, and other image processing on the first displayed content, such as user-specified uploaded pictures or videos, and finally generate a storyboard image that can be used as advertising material or as the main image for derivative advertising resources.
[0343] Similarly, the aforementioned "Wallpaper Inspiration," "Story Inspiration," and "Proposal Inspiration" can be controls for triggering wallpaper creation, story creation, and proposal creation, respectively, and will not be elaborated upon here.
[0344] S1104: Invoke the fourth generation model related to the fourth editing function and display the fifth editing interface.
[0345] For example, in response to the user operation detected in step S1103 above, mobile phone 100 can display the AI generation interface 020 corresponding to the "advertising storyboard" function shown in Figure 10b, i.e., the fifth editing interface. The fourth generation model related to the fourth editing function can be a generation model that mobile phone 100 has registered with the video editing capabilities or image processing functions of a large model and pre-built for processing target images or target videos such as storyboarding and creation. Based on the fourth generation model related to the fourth editing function, mobile phone 100 can also display the fifth editing interface including at least one parameter adjustment control. In some embodiments, when mobile phone 100 calls the fourth generation model, it can also load some pre-trained scene models, such as cross-country running character models, track models, skateboarding character models, track models, etc., and provide parameter adjustment controls on the displayed fourth editing interface for users to set relevant parameters, which is not limited here.
[0346] As shown in Figure 10b, the function option box 021 on the AI generation interface 020 can display the "Ad Storyboard" function that the user has selected. In some embodiments, the user can also click on the function option box 021 to reselect other functions, such as reselecting the aforementioned "Wallpaper Inspiration," "Story Inspiration," and "Proposal Inspiration" functions, without any restrictions.
[0347] It is understandable that the creative application running on the mobile phone 100 can directly generate creative storyboard images that meet the user's needs based on the user's settings of relevant parameters on the relevant interface, or it can generate creative storyboard images that meet the user's needs based on the user's uploaded images or videos and the user's settings of relevant parameters on the relevant interface, without any restrictions. For the latter, for example, the user can upload source images or videos for storyboard production in the AI generation interface 020 shown in Figure 10b, as the target image or video to be edited.
[0348] S1105: Detected user operation in the fifth editing interface to set control parameters for at least one parameter adjustment control.
[0349] For example, taking the AI generation interface 020 shown in Figure 10b as the fifth editing interface mentioned above, the AI generation interface 020 may include multiple parameter adjustment controls, such as the various parameter adjustment controls related to "posture adjustment," "character preset," and "venue preset" in the parameter adjustment control area 023. For example, the user can set the parameter related to "scene adjustment" for "posture adjustment" to "close-up" and the parameter related to "angle adjustment" to "side view." As another example, the user can set the parameter for "character preset" to "young male," and the user can also set the parameter for "venue preset" to "track and field." The "character preset" can set the age range of the character; in other embodiments, the "character preset" can also set the gender, occupation, etc., of the character, which is not limited here.
[0350] In other embodiments, the parameters that can be set for the above-mentioned parameter adjustment controls may also include other parameter options, such as the "long shot" parameter corresponding to "shot size adjustment", the "front" and "back" parameters corresponding to "angle adjustment", the "young man" parameter corresponding to "character preset", and the "swimming pool" and "gym" parameters corresponding to "location preset", etc., without limitation.
[0351] S1106: Display at least one prompt related to setting control parameters in the fifth editing interface.
[0352] For example, in response to the user's setting operations on various parameter adjustment controls in the parameter adjustment control area 023 shown in Figure 10b, the mobile phone 100 can display at least one prompt corresponding to each parameter in the description information input box 022 of the AI generation interface 020. This prompt may include, for example, "Shot adjustment: Close-up," "Angle adjustment: Side view," "Character preset: Young man," "Field preset: Track and field," etc. Furthermore, in some embodiments, the description information input box 022 may also support user input of description information or prompts.
[0353] S1107: The user has confirmed the start of the generation operation on the fifth editing interface.
[0354] For example, continuing to refer to Figure 10b, if the mobile phone 100 detects that the user clicks the "Start Generation" control 024 in the AI generation interface 020, it will detect that the user has confirmed the start of generation on the fifth editing interface mentioned above.
[0355] S1108: Generate and display the fifth processing result based on at least one prompt word.
[0356] For example, after the mobile phone 100 detects that the user has confirmed the start of the generation operation on the fifth editing interface in the above step S1107, it can generate a storyboard image corresponding to the prompt information generated according to the prompt information displayed in the description information input box 022, as the fifth processing result.
[0357] In other embodiments, if a user uploads target images or target videos for generating storyboard images through the AI generation interface 020, the mobile phone 100 can also generate corresponding storyboard images based on at least one of the prompt words and the target images or videos uploaded by the user.
[0358] As an example, the interface displaying the fifth processing result can be referenced to the generation result interface 030 shown in Figure 10c. This generation result interface 030a can display the storyboard image 030a generated based on at least one of the aforementioned prompts, etc. In some embodiments, the generation result interface 030 can also simultaneously display other content that may be generated based on at least one of the aforementioned prompts, etc., without limitation.
[0359] Referring again to 10c, the generated result interface 030 can also display application controls 031, copy controls 032, regenerate controls 033, and share controls 034. Among them, application control 031 can control the application of the currently displayed fifth processing result, i.e., storyboard image 030a, to the target application that calls the AI generation model, such as the creation application described above in the embodiments of this application.
[0360] S1109: A user instruction to apply the fifth processing result to the ninth interface was detected.
[0361] S1110: Displays the tenth interface, which includes the fifth processing result.
[0362] For example, if the mobile phone 100 detects a user's click operation on the application control 031 in the generated result interface 030, it can apply the generated storyboard image 030a to the creation application and display the gallery interface 040 shown in Figure 10d.
[0363] It is understandable that the description information and prompts displayed in the description information input box of the fifth editing interface, such as the description information input box 022 shown in Figure 10b, can all be converted into prompt information based on the prompt generation model integrated during the construction of the fourth generation model. This information is then input into the large model or AI model upon which the fourth generation model depends for model inference. At this point, the prompt generation model of the fourth generation model, based on the prompt information generated by the user through at least one of the aforementioned parameter tuning controls, may be more accurate than the prompt information manually entered by the user in the description information input box. Furthermore, combining the user's input in the description information input box with the parameters set by the user through at least one of the aforementioned parameter tuning controls can generate even more accurate prompt information. Thus, more accurate prompt information allows for more precise matching of the processing results to user needs, improving the user experience.
[0364] Furthermore, the processing results obtained by the mobile phone 100 based on the data processing method provided in this application can be directly applied to the relevant interface of the target application that calls the corresponding function based on the AI model. This simplifies the process for users to perform complex operations such as copying, pasting, saving, or loading the processing results in the AI application and then apply them to the target application, which also helps to improve the user experience.
[0365] The following section, with reference to the accompanying drawings, provides a detailed description of the operating system composition of the mobile phone and other electronic devices used in the data processing method provided in this application.
[0366] Figure 12a is a schematic diagram of the software structure of an operating system according to an embodiment of this application.
[0367] As shown in Figure 12a, the operating system on an electronic device, such as a mobile phone 100, can adopt a layered architecture. This layered architecture may include an application layer 310, an application framework layer 320, a system service layer 330, and a kernel layer 340.
[0368] The application layer 310 may include a series of installed applications, including system applications and third-party applications. System applications include memo apps and photo galleries, while third-party applications include image enhancement apps, creation apps, and AI applications developed based on large models. When running, apps such as memo apps, photo galleries, and creation apps can call the capabilities of various generative models provided by the large model used by the AI application to provide users with corresponding editing functions. These apps can be referred to as target applications in this embodiment. In some embodiments of this application, the system of electronic devices such as mobile phones 100 may not have AI applications installed; instead, some application programming interfaces (APIs) provided by the large model may be deployed in the application framework layer 320 and system service layer 330 to support calling the corresponding generative models. No restrictions are imposed here.
[0369] The application framework layer 320 can provide the application layer 310 with an capability framework 321, a user interface (UI) framework 322, and a user program framework 323, etc. In some embodiments, these frameworks may also be referred to as distributed frameworks. The UI framework 322 can integrate the capabilities provided by one or more generative models 22 provided by the large model to support the corresponding editing functions provided by the data processing method provided by the application layer 310 based on the data processing method provided in the embodiments of this application.
[0370] Referring to Figure 12b, the structure of each generation model 22 may include a prompt generation model 22a, a data association module 22b, a large model capability registration module 22c, and a parameter tuning module 22d. The prompt generation model 22a can generate prompt information based on control parameters set by the user through parameter tuning controls and user-input descriptive information. For example, the prompt generation model 22a can generate relatively accurate prompt information based on user input operations on the editing function interfaces provided by the generation models in the aforementioned target applications, operations performed on parameter tuning controls, and operations performed on association controls. The user's operations on association controls can trigger the acquisition of feature vectors or feature clustering based on user behavior data from the data association module 22b. It is understandable that the aforementioned prompt information can be input into the large model along with the target text, target image, target audio, or target video specified by the user. This prompt information can be used to control the model's inference process, and thus, when generating text, images, audio, video, and other content, it can incorporate user input operations on relevant interfaces or user behavior data, making the final generated content more in line with user needs.
[0371] The data association module 22b can extract feature vectors from local data and perform features clustering and other processing on the extracted feature vectors. It can then provide the extracted feature vectors and the results of features clustering and other processing to the prompt generation model 22a. The data association module 22b can associate data with the corresponding editing functions provided by the target application, such as associating family member descriptions with corresponding photos.
[0372] The large model capability registration module 22c can be used to register the relevant capabilities provided by large models for computation, inference, and content generation, such as the capabilities required for the aforementioned text expansion, poetry and prose, AI painting, and story creation functions. It can be understood that the generating model 22 can call the registered large model capabilities through the corresponding APIs. The large models providing these capabilities can include cloud-side models deployed on cloud servers, or edge-side models deployed on terminals with strong computing power; there are no restrictions here.
[0373] The parameter tuning module 22d can be used to adjust the relevant parameters of the large model capabilities called based on the user's settings on the editing function interface provided by the target application. For example, if the user sets "word count", "tone", "stage" on the writing assistant interface, the corresponding parameters of the large model capabilities called can be adjusted to control the generation model to call the corresponding large model capabilities to generate content that conforms to the user's settings.
[0374] The system service layer 330 can provide services to the applications and SDKs in the application layer 310 through the application framework layer 320. In this embodiment, the system service layer 330 may include database services such as a feature vector database 331, used to store and manage feature vectors extracted from local data, and to obtain feature clustering results based on feature vector clustering analysis. Local data may include operation data generated by the user while using electronic devices such as mobile phones 100, also known as user behavior data. The generative model of the application framework layer 320 can generate corresponding prompt information based on the feature vectors and feature clustering results provided by the feature vector database 331 and other local data processing results, as input to the large model for model inference. This allows for intelligent recommendations by integrating user behavior data when responding to calls from target applications, thereby improving the user experience of the target application and its functions.
[0375] The application framework layer 320 and the system service layer 330 may also include a multimodal input subsystem 332 and an AI subsystem 333, etc. Among them, the multimodal input subsystem 332 can support processing multimodal data input by users, such as text, images, voice, video, etc., into input data that can be recognized by other software structures of the system, and then provide it to other subsystems or frameworks, system services, such as the AI subsystem 333, for further processing.
[0376] The AI subsystem can integrate the APIs for calling various generative models provided by the large model, as well as the input and output interfaces of the large model, providing system services such as a capability registration framework and a capability execution framework for the application layer 310. In this embodiment, the capabilities (or services) required by the target application to provide corresponding editing functions can be pre-registered in the AI subsystem 333. Furthermore, the corresponding loading modules provided by the AI subsystem 333, such as the capability configuration loading module, configure the relevant parameters of the generative model and load the corresponding generative model to provide the corresponding registered capabilities.
[0377] As an example, referring to Figure 12c, in some embodiments of this application, the AI subsystem 333 may include a capability registration framework and a capability execution framework. The capability registration framework may include a cloud-side model capability registration module 33a and an edge-side model capability registration module 33b. The cloud-side model capability registration module 33a can be used to define standard APIs and support the registration of parameter templates, large model capabilities, and other related APIs provided by the cloud-side model. It can be understood that when the cloud-side model completes development or undergoes a version update, the cloud-side model capability registration module 33a can support loading the development data package or update data package uploaded by the corresponding integrated development environment (IDE) and loading or updating the relevant capabilities provided by the registered cloud-side model. In this way, the relevant capabilities of the registered cloud-side model can be provided without installing an AI application on electronic devices such as mobile phones.
[0378] Similarly, the edge model capability registration module 33b can be used to define standard APIs and support the registration of parameter templates, large model capabilities, and other related APIs provided by edge models. In addition, the edge model capability registration module 33b can preset a whitelist of target applications that support calling edge model capabilities or other types of whitelists. The registration of edge model capabilities can be controlled through this whitelist. No restrictions are imposed here, nor will they be elaborated upon.
[0379] Referring again to Figure 12c, the capability execution framework of the AI subsystem 333 may include a capability configuration loading module 33c, a cloud-side model capability invocation module 33d, and an edge-side model capability invocation module 33e.
[0380] The capability configuration loading module 33c can be used to load the relevant model components of the large model capabilities registered by the cloud-side model capability registration module 33a and the edge-side model capability registration module 33b. These model components may include neural network algorithms or deep learning models with configurable parameters, etc., without limitation. After loading the relevant model components, the capability configuration loading module 33c can configure the relevant parameters of the model components to control whether they can provide the corresponding large model capabilities when called and executed. In addition, the capability configuration loading module 33c can also be used for list authentication, such as authenticating the target application that will call the corresponding large model capabilities based on the whitelist preset by the edge-side model capability registration module 33b, etc.
[0381] The cloud-side model capability invocation module 33d can be used to invoke the capabilities of large models registered with the cloud-side model to support the corresponding generated models in implementing the corresponding editing functions in the target application. The cloud-side model capability invocation module 33d, in conjunction with the aforementioned capability configuration loading module 33c, can control the construction of the relevant parameter adjustment page. The cloud-side model capability invocation module 33d can also construct a prompt generated model 22a for the corresponding generated model 22, thus serving as a component of the software structure of that generated model 22.
[0382] The edge-side model capability invocation module 33e can be used to invoke the large model capabilities registered with the edge-side model to support the corresponding generation model in implementing the corresponding editing functions in the target application. The edge-side model capability invocation module 33e, in conjunction with the aforementioned capability configuration loading module 33c, can control the construction of the relevant parameter adjustment page and can perform complex parameter settings or adjust page loading. These complex parameters may include, for example, prompts generated based on data related to human characteristics in the local feature database, text recorded by memo applications, and descriptive information of images. The data related to human characteristics may be feature data extracted from images or videos obtained from gallery applications and stored in the local feature database. This local feature database can include not only the aforementioned human characteristics but also other features, such as feature data describing the body shape, expression, and clothing characteristics of the person indicated by the human characteristics; there are no restrictions on this. The edge-side model capability invocation module 33e can also construct a prompt generation model 22a for the corresponding generation model 22, thus serving as a component of the software structure of the generation model 22.
[0383] It is understandable that, for the same generative model, in order to achieve the editing functions provided by the target application, the capabilities of the generative model can be composed by registering and loading one or more large model capabilities of the cloud-side model, or by registering and loading one or more large model capabilities of the edge-side model, or by registering and loading large model capabilities of the cloud-side model and the edge-side model respectively. No restrictions are imposed here.
[0384] Kernel layer 340 is the layer between hardware and software. The kernel layer 340 of the operating system shown in Figure 3 may include a kernel subsystem and a driver subsystem. The kernel subsystem, given that distributed operating systems can employ a multi-kernel design, supports the selection of a suitable operating system (OS) kernel for different resource-constrained devices.
[0385] The kernel abstract layer (KAL) on the kernel subsystem provides basic kernel capabilities to the upper layers by shielding the differences between multiple kernels, including process / thread management, memory management, file system, network management, and peripheral management.
[0386] The driver subsystem provides a driver framework that forms the foundation for the open hardware ecosystem of some distributed systems, offering unified peripheral access capabilities and a framework for driver development and management. The kernel layer includes at least display drivers, camera drivers, audio drivers, and sensor drivers.
[0387] Based on the operating system structure shown in Figure 12a above, Figure 12d illustrates a schematic diagram of the implementation principle of generating expanded content by incorporating associated data, according to an embodiment of this application.
[0388] As shown in Figure 12d, the implementation principle includes the following process:
[0389] 1201: Constructing a local feature database. This specifically includes: extracting feature vectors from local data and constructing a feature vector database; performing feature clustering on the feature vectors in the database to determine the relationships between them. For example, for people in locally stored images or videos managed by a photo library application, the extracted feature vectors can include image features of multiple different characters, including facial features, action features, posture features, clothing features, age features, and gender features. Based on this, for different characters, these image features can be clustered to form feature vector databases for each character, such as a feature vector database for "father," a feature vector database for "mother," and a feature vector database for "younger brother," thus constructing a local feature database.
[0390] It is understandable that the naming of the aforementioned characters such as "father," "mother," and "brother" can be done in advance by extracting keywords and marking them based on the character descriptions or notes entered by the user, so as to be used for related data retrieval in the following text.
[0391] 1202: Keyword Extraction from Input Information. This specifically includes performing natural language processing (NLP) semantic understanding on user input and extracting keywords. This input information can include descriptive information or notes entered by the user on photos taken in the relevant editing interface of a photo gallery application, such as "father," "mother," "baby," or notes about travel destinations. It can also include personal size information entered by the user in other applications, such as shopping applications, etc., without limitation.
[0392] Electronic devices like Mobile 100 can use NLP technology to extract keywords from various user input information, which can then be used as search keywords for relevant data. For example, if a user enters "dad" or clicks on a related category marked with "dad," Mobile 100 or similar devices can extract the keyword "dad." Similarly, if a user enters "brother" or clicks on a related category marked with "brother," Mobile 100 or similar devices can extract the keyword "brother."
[0393] 1203: Feature Data Association. This specifically includes: based on the keywords extracted above, performing associated data retrieval in conjunction with a feature vector database, and then associating the keywords with the feature vectors. For example, based on the extracted keyword "father," the feature vector database for "father" can be retrieved for data association; based on the extracted keyword "younger brother," the feature vector database for "younger brother" can be retrieved for data association.
[0394] This feature vector database can store feature vectors corresponding to relevant content in images, audio, or video associated with corresponding keywords. This association can be pre-established based on local data generated by users using mobile phones and other electronic devices, through feature extraction, feature clustering, and classification storage, which will not be elaborated here.
[0395] 1204: Integrating related data to generate prompt information. Specifically, this includes: a function panel displaying parameter tuning controls, supporting user selection of related information, and then requesting feature information from the local feature database, such as obtaining the aforementioned feature vectors or feature clustering results. This function panel can, for example, include the display panel of the writing assistant interface 430 shown in Figure 4c of Embodiment 2. Correspondingly, the local feature database can recommend related data and present it on the function panel, such as the image gallery character-related content provided in Embodiment 2. Based on this, the parameter tuning results are integrated with feature information, combined with preset feature information weights or parameter tuning result weights, to generate prompt information. For example, the prompt information generated for the character "father" may include "father," "tall," and "thin," while the prompt information generated for the character "younger brother" may include "younger brother," "face," and "colorful." The generated prompt information is used for model inference, which can control the corresponding generation model to obtain more accurate processing results that better meet user expectations or user needs, including the aforementioned first to fifth processing results.
[0396] In other embodiments, the aforementioned associated data may include not only image data and / or video data provided by a gallery application, but also audio data provided by a recording application, calendar data provided by a calendar application, map data provided by a map application, email data provided by an email application, text data provided by a notepad or other text editing application, contact data provided by a contact application, and SMS data provided by an SMS application, etc., without limitation.
[0397] Figure 13 shows a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application.
[0398] Mobile phone 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, buttons 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identity module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0399] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0400] Processor 110 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0401] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0402] In this embodiment, the processor 110 of the mobile phone 100 can generate operation control signals through the controller to control the fetching and execution of instructions related to each step in the implementation process provided in the above embodiments, so as to realize the data processing method provided in this application.
[0403] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the aforementioned memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0404] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a universal serial bus (USB) interface, etc.
[0405] USB port 130 is a USB standard compliant interface, which may include Mini USB, Micro USB, USB Type-C, etc. USB port 130 can be used to connect a charger to charge mobile phone 100, and can also be used for data transfer between mobile phone 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0406] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may also adopt different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0407] The charging management module 140 receives charging input from a charger. The charger may include a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the mobile phone 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0408] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0409] The wireless communication function of mobile phone 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.
[0410] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0411] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the mobile phone 100.
[0412] The wireless communication module 160 can provide solutions for wireless communication applications on the mobile phone 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc.
[0413] In some embodiments, antenna 1 of mobile phone 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling mobile phone 100 to communicate with networks and other devices via wireless communication technology. The aforementioned wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The aforementioned GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0414] The mobile phone 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0415] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini-LED, Micro-LED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc. In some embodiments, the mobile phone 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0416] The mobile phone 100 can achieve shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0417] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits this electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0418] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element may include a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0419] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when a mobile phone 100 is selecting a frequency, the DSP performs Fourier transforms on the frequency energy.
[0420] Video codecs are used to compress or decompress digital video. Mobile phone 100 can support one or more video codecs. Thus, mobile phone 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0421] An NPU (Neural Processing Unit) is a neural network (NN) computing processor that, by borrowing from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, rapidly processes input information and can continuously learn on its own. The NPU enables intelligent cognitive applications on the mobile phone 100, such as image recognition, face recognition, speech recognition, and text understanding. In this embodiment, the mobile phone 100 can use the NPU to pre-build generative models with various editing functions. The capabilities of the large model upon which each generative model relies to achieve its corresponding editing function can be obtained by registering the corresponding API from the large model.
[0422] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0423] Internal memory 121 can be used to store computer executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of mobile phone 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of mobile phone 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.
[0424] The mobile phone 100 can achieve audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0425] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A may be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Mobile phone 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, mobile phone 100 detects the intensity of the touch operation based on pressure sensor 180A. Mobile phone 100 may also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities may correspond to different operation commands.
[0426] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of mobile phone 100, in a different position than display screen 194.
[0427] Keypad 190 includes a power button, volume buttons, etc. Keypad 190 may include mechanical keys or touch keys. Mobile phone 100 can receive keypad input and generate key signal inputs related to user settings and function control of mobile phone 100.
[0428] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0429] Indicator 192 may include indicator lights, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0430] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the mobile phone 100. In some embodiments, the mobile phone 100 may also use an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the mobile phone 100 and cannot be separated from it.
[0431] This application also provides a computer program product for implementing the data processing methods provided in the above embodiments.
[0432] Various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or combinations of these implementation methods. Embodiments of this application can be implemented as computer program modules or module code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0433] Computer program modules or module code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.
[0434] Module code can be implemented using a high-level modular language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used to implement module code when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can include compiled or interpreted languages.
[0435] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, optical discs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0436] In this specification, the reference to "an embodiment" or "an embodiment" means that a specific feature, structure, or characteristic described in connection with the embodiment is included in at least one exemplary implementation or technology disclosed according to an embodiment of this application. The appearance of the phrase "in an embodiment" in various places in the specification does not necessarily refer to the same embodiment.
[0437] The disclosure of embodiments of this application also relates to means for performing operations in text. This means may be specifically constructed for the claimed purpose or may include a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored on a computer-readable medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, application-specific integrated circuits (ASICs), or any type of medium suitable for storing electronic instructions, and each may be coupled to a computer system bus. Furthermore, the computer mentioned in the specification may include a single processor or may include an architecture employing multiple processors for increased computing power.
[0438] Furthermore, the language used in this specification has been primarily chosen for readability and instructional purposes and may not have been chosen to depict or limit the disclosed subject matter. Therefore, the embodiments disclosed herein are intended to illustrate, and not limit, the scope of the concepts discussed herein.
Claims
1. A data processing method applied to an electronic device, characterized in that, The method comprises: displaying a first interface of a first application, wherein the first interface comprises first display content; detecting a first processing instruction for the first display content, processing the first display content according to reference data, and generating a first processing result, wherein the first processing result comprises information obtained according to the reference data; displaying the first processing result in the first interface.
2. The method of claim 1, wherein, The reference data is data stored by the electronic device.
3. The method of claim 2, wherein, The reference data comprises one or more of the following data: picture data; video data; audio data; schedule data; map data; email data; text data; contact data; and short message data.
4. The method according to any one of claims 1 to 3, characterized in that, The reference data comprises application data of a second application, wherein the application data of the second application comprises data generated during execution of the second application by the electronic device and / or data obtained by the second application from a local storage space of the electronic device.
5. The method of claim 4, wherein, The second application comprises one or more of the following applications: a gallery application; a recording application; a schedule application; a map application; an email application; a notepad application; a text editing application; a contact application; and a short message application.
6. The method of claim 3, wherein, In a case where the reference data comprises the picture data and / or the video data, the information obtained according to the reference data comprises one or more of the following: characteristic information of one or more characters in the picture data and / or the video data; character relationship information of multiple characters in the picture data and / or the video data; environmental characteristic information of a background of one or more characters in the picture data and / or the video data.
7. The method of claim 6, wherein: The character characteristic information comprises one or more of the following characteristic information of a character: appearance characteristic, action characteristic, clothing characteristic, age characteristic, and gender characteristic.
8. The method according to any one of claims 1 to 7, characterized in that, The processing of the first display content according to the reference data to generate the first processing result comprises: obtaining at least one prompt word according to the reference data, and processing the first display content based on the at least one prompt word to generate the first processing result.
9. The method of claim 8, wherein, The obtaining of the at least one prompt word according to the reference data comprises: determining at least one control parameter according to the reference data, wherein the control parameter comprises the reference data or a description of the reference data; performing conversion processing on the at least one control parameter to obtain the at least one prompt word.
10. The method of claim 9, wherein, The electronic device comprises a first model, and The processing of the first display content based on the at least one prompt word to generate the first processing result comprises: calling the first model, and inputting the first display content and the at least one prompt word into the first model for model inference to generate the first processing result.
11. The method according to claim 9 or 10, characterized in that, The at least one control parameter further comprises a parameter determined by a user's parameter adjustment operation on at least one parameter adjustment control.
12. The method of claim 11, wherein, In a case where the first processing result comprises text, the at least one control parameter comprises one or more of the following: a word count for limiting the length of the text content; a tone for limiting the description style of the text content; and a target audience category for limiting the text content.
13. The method of claim 11, wherein, In a case where the first processing result comprises a picture or a video, the at least one control parameter comprises one or more of the following: a style parameter for defining a picture style or a video style; a size parameter and / or a material parameter for defining a dress of a character subject in the picture or the video; a scene parameter and an angle parameter for defining a highlight effect of the character subject in the picture or the video; a parameter for defining an age of the character in the picture or the video; a parameter for defining a location where the character is in the picture or the video.
14. The method according to any one of claims 1 to 13, characterized in that, Before processing the first display content according to the reference data, the method further comprises: detecting a first processing instruction for the first display content, and displaying a second interface, wherein the second interface comprises an association control for triggering acquisition of the reference data required for generating the first processing result.
15. The method of claim 14, wherein, Before processing the first display content according to the reference data, the method further comprises: detecting that the association control is placed in an open state, and acquiring the reference data.
16. The method of claim 15, wherein, The detecting that the association control is placed in the open state and acquiring the reference data comprises: displaying selection controls corresponding to at least one classification data related to the reference data, wherein the at least one classification data comprises a first classification data corresponding to a first selection control; detecting a first operation of the user selecting the first selection control, and acquiring the first classification data as the reference data.
17. The method of claim 14, wherein, The second interface further comprises at least one parameter adjustment control, and The method further comprises: detecting a parameter adjustment operation of the user on the at least one parameter adjustment control, and processing the first display content according to the reference data and a parameter determined based on the parameter adjustment operation, to generate a first processing result.
18. The method of claim 14, wherein, The second interface further comprises a generation control for indicating start of generation of the processing result, and The method further comprises: detecting a second operation of the user clicking the generation control, and displaying a third interface, wherein the third interface comprises the first processing result obtained by processing the first display content according to the reference data.
19. The method of claim 18, wherein, The third interface further comprises an application control for applying the first processing result to the first interface, and The displaying of the first processing result in the first interface comprises: detecting a third operation of the user clicking the application control, and displaying the first processing result in the first interface.
20. The method of any one of claims 1 to 19, wherein, In a case where the first application is a memo application and the first display content comprises text or a picture, the detecting of the first processing instruction for the first display content comprises any one of the following: detecting the first processing instruction based on a fourth operation of the user long-pressing the text in the first interface and selecting the first display content; detecting the first processing instruction based on a fifth operation of the user long-pressing the first display content.
21. The method of any one of claims 1 to 19, wherein, In a case where the first application is any one of a gallery application, an image beautification application, and a creation application, and the first display content comprises a picture or a video, the detecting of the first processing instruction for the first display content comprises any one of the following: The first processing instruction is detected based on a sixth operation of a user long-pressing the first display content; The first processing instruction is detected based on a seventh operation of a user clicking a processing control on the first interface, wherein the processing control on the first interface includes a control indicating processing of the first display content on the first interface.
22. An electronic device, comprising: Comprise: One or more processors; One or more memories; the one or more memories store one or more programs, when the one or more programs are executed by the one or more processors, make the electronic device execute the data processing method of any one of claims 1 to 21.
23. A computer readable medium characterized by The readable medium stores instructions, which when executed on a computer, cause the computer to execute the data processing method of any one of claims 1 to 21.
24. A computer program product, characterised in that, Comprise computer programs / instructions, which when executed by a processor, implement the data processing method of any one of claims 1 to 21.
Citation Information
Patent Citations
Multimedia content generation method and device, equipment and storage medium
CN116701669A
Human-computer interaction method, device and equipment and storage medium
CN117032515A
Image transformation method, electronic equipment and storage medium
CN117170560A
Human-computer interaction method, display method, apparatus, and device
WO2024061163A1