Content processing method and system and electronic equipment

By using the context information of the content in the AI ​​model for task processing, the instability problem of the AI ​​model in complex task scenarios is solved, and the accuracy and practicality of task processing are improved.

CN120146004APending Publication Date: 2025-06-13HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311705791.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing AI models have unstable task execution effects, especially in complex task scenarios, which are difficult to meet task processing requirements, resulting in low practicality.

Method used

By obtaining content, location information and task information, inputting them into the AI ​​model, and using the context information in the content for task processing, thereby improving the accuracy of processing results.

Benefits of technology

The effect and practicality of the AI ​​model in task processing is improved, making the output results easier to meet task processing requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146004A_ABST
    Figure CN120146004A_ABST
Patent Text Reader

Abstract

The invention provides a content processing method and system and electronic equipment, and the method comprises the steps: obtaining first content, first position information used for indicating a first position in the first content, and first task information used for indicating the execution of a first processing task for the first position by the electronic equipment; and inputting the first content, the first position information and the first task information into a first model, so that the first model executes a first processing task for the first position according to a part of or all of the content except the first position in the first content, and then second content is obtained. Wherein part or all of the first content except the first position can be used as the context content of the first position, and the electronic device performs task processing by using the first model in combination with the context content, so that the accuracy of an output result when the first model is used for executing the task is improved; and the output result can more easily meet the task processing requirement, so that the practicability of the first model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of electronic devices, and in particular, to a content processing method, system, and electronic device. Background Art

[0002] An artificial intelligence (AI) model (or referred to as a large AI model or a large pre-trained AI model) is a model that can adapt to a series of tasks after being trained with a large amount of data. The AI model belongs to a machine learning model. Currently, some AI models (such as generative AI models, large language models, etc.) can execute corresponding tasks based on the input task indication information and output task processing results. Alternatively, the AI model can execute corresponding tasks on the input content and output task processing results based on the input content and the task indication information corresponding to the content. However, in actual applications, the task execution effect of the AI model may be unstable, and the task processing results output by the AI model in some scenarios (such as complex task scenarios, etc.) may be difficult to meet the task processing requirements, resulting in low practicality of the AI model. Summary of the Invention

[0003] This application provides a content processing method, system, and electronic device to improve the effect of task processing by the model, thereby improving the practicality of the model.

[0004] In a first aspect, an embodiment of this application provides a content processing method, and the method includes: obtaining first content, first position information, and first task information; where the first position information is used to indicate a first position in the first content, and the first task information is used to indicate a first processing task to be executed for the first position; inputting the first content, the first position information, and the first task information into a first model to determine second content; where the second content includes: content obtained by the first model executing the first processing task for the first position according to first reference content; the first reference content includes part or all of the content in the first content except the first position. Optionally, the first model belongs to an AI model.

[0005] In this method, in the first content, the first reference content can serve as the context content of the first position. By inputting the first content, the first position information, and the first task information into the first model, the electronic device can enable the first model to refer to the context content of the first position to process the first position when performing the first processing task indicated by the first task information for the first position indicated by the first position information, thereby obtaining a more accurate processing result. Therefore, the above method helps to improve the accuracy of the result output when the electronic device uses the first model to perform tasks, making the output result more likely to meet the task processing requirements, and further improving the practicality of the first model. When the first model is an AI model, the above method can improve the effect of processing tasks through the AI model, and further improve the practicality of the AI model.

[0006] In a possible design, the first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content. In this method, the context content of the first position can be the content of the position on either side of the first position, with high flexibility and facilitating the control of the data volume of the context content of the first position, thereby ensuring a certain processing efficiency.

[0007] Optionally, if the first position is a position before the first sub-content, then in the content input to the first model, the first position information and the first task information are before the first sub-content; if the first position is a position after the first sub-content, then in the content input to the first model, the first position information and the first task information are after the first sub-content. In this method, the positional relationship between the position information / task information and the sub-content in the content input to the first model can be associated with the positional relationship between the position corresponding to the position information and the sub-content in the first content, which can improve the consistency before and after in the content conversion process. At the same time, if the content input to the first model is displayed, it can improve the readability of the content input to the first model.

[0008] In a possible design, the first content further includes a second sub-content, the first position is a position between the first sub-content and the second sub-content, and the first reference content further includes the second sub-content. In this method, the context content of the first position can be the content of the positions on both sides of the first position, with high comprehensiveness and helping to improve the accuracy of task processing.

[0009] Optionally, in the content input to the first model, the first position information and the first task information are between the first sub - content and the second sub - content. In this method, the positional relationship between the position information / task information and the sub - content in the content input to the first model can be associated with the positional relationship between the position corresponding to the position information and the sub - content in the first content, which can improve the consistency before and after in the content conversion process. At the same time, if the content input to the first model is displayed, it can improve the readability of the content input to the first model.

[0010] In a possible design, the first content includes a third sub - content, and the third sub - content is located at the first position; the second content specifically includes the content obtained by the first model performing the first processing task on the third sub - content according to the first reference content.

[0011] Combined with the above method, there may or may not be a third sub - content at the first position in the first content, that is, the content may be empty. Whether there is a sub - content at the first position in the first content or not, the electronic device can, based on the above method, combine the context content of the first position and perform relatively accurate task processing on the first position. Therefore, the above method has a good effect in task processing through the first model.

[0012] In a possible design, the first position information includes start position information and end position information; wherein, the start position information is used to indicate the starting point of the first position, and the end position information is used to indicate the ending point of the first position; the third sub - content is located between the starting point and the ending point in the first content.

[0013] In this method, the start position information and the end position information can clearly and accurately indicate the position of the third sub - content in the first sub - content, which is convenient for the first model to identify the third sub - content and perform subsequent processing, thereby improving the processing efficiency.

[0014] In a possible design, in the content input to the first model, the third sub - content is between the start position information and the end position information; and / or, in the content input to the first model, the first task information is before the start position information, or between the start position information and the third sub - content, or between the third sub - content and the end position information, or after the end position information.

[0015] In this method, the positional relationship between the start position information and the end position information in the position information and the sub - content, task information, etc. in the content input to the first model can be flexibly set, which can improve the flexibility and practicability of the solution.

[0016] In a possible design, the method is applied to a server; the obtaining of the first content, the first location information, and the first task information includes: receiving the first content, the first location information, and the first task information from a terminal device.

[0017] When the content processing method provided above is applied to a server, the server can obtain the content to be input into the first model, i.e., the first content, the first location information, and the first task information, from a terminal device, and use the first model to process this content. Therefore, the above method can be applied to the scenario where the server uses the first model to perform task processing on the content provided by the terminal device, thereby avoiding the workload of task processing on the terminal device and improving the processing efficiency on the terminal device side.

[0018] In a possible design, the obtaining of the first content, the first location information, and the first task information includes: obtaining the first content and the first indication information; wherein, the first indication information is used to indicate to perform the first processing task on the first location in the first content; and determining the first location information and the first task information according to the first content and the first indication information.

[0019] Among them, the first indication information can be used for an electronic device to identify or determine the first location and the first processing task, and the first location information and the first task information can be respectively used for the first model to identify or determine the first location and the first processing task. By determining the first location information and the first task information according to the first content and the first indication information, it is possible to convert the location information and task information recognizable by the electronic device into the location information and task information recognizable by the first model, so as to facilitate the first model to perform corresponding processing, thereby improving the processing efficiency and the accuracy of task execution of the first model.

[0020] In a possible design, after obtaining the first content and the first indication information and before determining the first location information and the first task information according to the first content and the first indication information, the method further includes: displaying the first content and the first indication information; receiving a first operation; wherein, the first operation is used to indicate: converting the first indication information into the first location information and the first task information according to the first content.

[0021] In this method, the electronic device can trigger the conversion of the execution indication information, i.e., the determination of the location information and the task information, based on the operation performed by the user, which facilitates the user to control the content processing process.

[0022] In a possible design, the obtaining of the first content and the first indication information includes: obtaining and displaying the first content through a first application, and determining the first indication information according to a second operation acting on the first content received; wherein, the second operation is used to indicate the first indication information; or, obtaining the first content and the first indication information indicated by a third operation received through the first application; or, receiving the first content and the first indication information sent by a second application through the first application; wherein, the second application is used to generate the first content and the first indication information.

[0023] Optionally, before receiving the first content and the first indication information sent by the second application through the first application, the method further includes: obtaining and displaying the first content through the second application, and determining the first indication information according to a fourth operation acting on the first content received.

[0024] Wherein, the first application and the second application are applications installed in the electronic device. The electronic device can generate and display the first content through the first application, and the user can edit the first content displayed by the first application by performing a second operation, and the first application of the electronic device can then determine the first indication information based on the second operation. Or, the user can directly indicate the first content and the second indication information to the first application of the electronic device by performing a third operation such as a paste operation. Or, the electronic device can generate the first content through the second application and determine the first indication information according to the operation performed by the user on the first content, and can send the first content and the first indication information to the first application through the second application. In the above method, the electronic device can use at least one application to obtain the first content and the first indication information in different ways, with relatively high flexibility and practicability.

[0025] Optionally, the inputting of the first content, the first location information, and the first task information into a first model to determine a second content includes: inputting the first content, the first location information, and the first task information into the first model through the first application to determine the second content. Optionally, before inputting the first content, the first location information, and the first task information into the first model through the first application to determine the second content, the method further includes: displaying the first content, the first location information, and the first task information through the first application; receiving a fifth operation through the first application; wherein, the fifth operation is used to indicate to perform a processing task according to the first content, the first location information, and the first task information.

[0026] Among them, by displaying the content to be input into the first model, namely the first content, the first position information, and the first task information, it enables the user to conveniently view the content input into the first model, and can trigger the execution of the content processing task based on the corresponding operations performed by the user. Therefore, it is convenient for the user to control the content processing process, thereby improving the observability and controllability of the content processing process.

[0027] In a possible design, the method is applied to a terminal device.

[0028] In a possible design, before inputting the first content, the first position information, and the first task information into the first model to determine the second content, the method further includes: obtaining second position information and second task information; wherein, the second position information is used to indicate a second position in the first content, and the second task information is used to indicate to perform a second processing task on the second position; when inputting the first content, the first position information, and the first task information into the first model, the method further includes: inputting the second position information and the second task information into the first model; the second content further includes: the content obtained by the first model performing the second processing task on the second position according to the second reference content; wherein, the second reference content includes part or all of the content in the first content except the second position.

[0029] Based on the above method, the electronic device can utilize the first model to simultaneously process multiple sets of tasks for the first content, such as the first processing task and the second processing task, thereby achieving the effect of simultaneously processing multiple tasks using the same model. Therefore, the above method can be applied to scenarios of simultaneously processing multiple tasks, and can improve the multi-task processing efficiency while ensuring the processing effects of each task in the multiple tasks.

[0030] In a possible design, before inputting the first content, the first position information, and the first task information into the first model to determine the second content, the method further includes: obtaining training data; wherein, the training data includes third content, third position information, third task information, and fourth content; wherein, the third position information is used to indicate a third position in the third content, and the third task information is used to indicate to perform a third processing task on the third position; the fourth content includes: the content obtained by performing the third processing task on the third position according to part or all of the content in the third content except the third position; training the set model according to the training data to obtain the first model.

[0031] Based on the above method, the electronic device can achieve the training of the first model for processing single tasks.

[0032] In a possible design, the obtaining of the training data includes: obtaining the third content and the second indication information; wherein, the second indication information is used to indicate to perform the third processing task on the third position in the third content; determining the third position information and the third task information according to the second indication information; inputting the third content and the second indication information into a second model to determine a fifth content; or, inputting the content at the third position in the third content and the second indication information into the second model to determine a fifth content; or, when there is no content at the third position, inputting the second indication information into the second model to determine a fifth content; wherein, the fifth content includes the content obtained by the second model performing the third processing task according to the input content; determining the fourth content according to the fifth content.

[0033] Optionally, the determining the fourth content according to the fifth content includes: using the fifth content as the fourth content, or, in response to a received sixth operation, modifying the fifth content to obtain the fourth content; wherein, the sixth operation is used to indicate modifying the fifth content to the fourth content. Wherein, the second model may be a set generative AI model.

[0034] The above method for obtaining training data can be applied to the scenario of training a first model for processing single tasks. Through the above method, the electronic device can use a set second model to realize the acquisition of partial training data used as the task processing result. Among them, by modifying the content obtained by the second model processing, that is, the fifth content, according to the user operation, the accuracy of the task processing result data in the training data can be improved by means of manual correction. Further, the accuracy of model training using the training data can be improved, and then a first model with better processing effect can be obtained.

[0035] In a possible design, the training data further includes fourth position information and fourth task information; wherein, the fourth position information is used to indicate a fourth position in the third content, and the fourth task information is used to indicate to perform a fourth processing task on the fourth position; the fourth content further includes: the content obtained by performing the fourth processing task on the fourth position according to part or all of the content in the third content except the fourth position.

[0036] Based on the above method, the electronic device can realize the training of a first model for unified processing of multiple tasks.

[0037] In a possible design, the obtaining of the training data includes: obtaining the third content, the second indication information, and the third indication information; wherein, the second indication information is used to indicate to perform the third processing task on the third position in the third content; the third indication information is used to indicate to perform the fourth processing task on the fourth position in the third content; inputting the third content and the second indication information into a second model to determine a fifth content; or, inputting the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, inputting the second indication information into the second model to determine the fifth content; wherein, the fifth content includes the content obtained by the second model performing the third processing task according to the input content; inputting the third content and the third indication information into the second model to determine a sixth content; or, inputting the content at the fourth position in the third content and the third indication information into the second model to determine the sixth content; or, inputting the third indication information into the second model to determine the sixth content; wherein, the sixth content includes the content obtained by the second model performing the fourth processing task according to the input content; determining the fourth content according to the fifth content and the sixth content.

[0038] Optionally, the determining the fourth content according to the fifth content and the sixth content includes: using the content obtained by combining the seventh content determined according to the fifth content and the eighth content determined according to the sixth content as the fourth content; wherein, the seventh content is the fifth content or the content obtained by modifying the fifth content in response to a received operation; the eighth content is the sixth content or the content obtained by modifying the sixth content in response to a received operation.

[0039] The above method for obtaining training data can be applied to the scenario of training a first model for processing multiple tasks. Through the above method, an electronic device can use a set second model to achieve the acquisition of partial training data for use as task processing results. Among them, by modifying the content obtained by the second model processing, namely the fifth content and the sixth content, according to user operations, the accuracy of the task processing result data in the training data can be improved in an artificial correction manner. Further, the accuracy of model training using the training data can be improved, and thus a first model with better processing effects can be obtained.

[0040] In a possible design, training the set model according to the training data to obtain the first model includes: inputting the training data into the set model to determine a ninth content; wherein, the ninth content includes the content obtained by the set model performing the third processing task on the third position in the third content; adjusting the set model according to the third content and the ninth content to obtain the first model.

[0041] In the above method, the electronic device can adjust the set model according to the deviation between the task processing result with higher accuracy, i.e., the ninth content, and the task processing result obtained by the set model, i.e., the third content, so that the task processing result of the set model approaches the more accurate task processing result, and then train to obtain a first model with better processing effect.

[0042] In a possible design, the ninth content further includes the content obtained by the set model performing the fourth processing task on the fourth position in the third content. Combining the above method, a first model with better processing effect and used for processing multiple tasks can be trained.

[0043] Optionally, the set model is a model with a set network structure or a set generative AI model.

[0044] In a possible design, the first content includes at least one of the following: text, image, audio, video, code.

[0045] The first content in the above method can include at least one type of content. Therefore, based on the above method, single-task or multi-task processing for various types of content can be realized, and better task processing results can be obtained.

[0046] In a possible design, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.

[0047] Among them, the semantic mask image can clearly and intuitively represent the positions of different types of content in the image, which helps to improve the accuracy of position recognition and task processing for the image.

[0048] In a possible design, the first model belongs to a generative artificial intelligence (AI) model. Among them, the generative AI model can perform generative tasks. By applying the generative AI model to the above method, the range of tasks that the model can execute can be further expanded, and the task execution ability of the model can be improved.

[0049] In a second aspect, the present application provides a content processing system, which includes a terminal device and a server; the terminal device is configured to: send a first content, first location information, and first task information to the server; wherein, the first location information is used to indicate a first location in the first content, and the first task information is used to indicate to perform a first processing task on the first location; the server is configured to: receive the first content, the first location information, and the first task information from the terminal device; input the first content, the first location information, and the first task information into a first model to determine a second content, wherein the second content includes content obtained by the first model performing the first processing task on the first location according to a first reference content, and the first reference content includes part or all of the content in the first content except the first location; send the second content to the terminal device; the terminal device is further configured to: receive the second content from the server.

[0050] In this method, in the first content, the first reference content can serve as the context content of the first location. By inputting the first content, the first location information, and the first task information into the first model, the server can enable the first model to refer to the context content of the first location to perform task processing on the first location when performing the first processing task indicated by the first task information for the first location indicated by the first location information, thereby obtaining a more accurate processing result. Therefore, the above method helps to improve the accuracy of the result output when the server uses the first model to perform tasks, making the output result easier to meet the task processing requirements, and further improving the practicality of the first model. When the first model is an AI model, the above method can improve the effect of processing tasks through the AI model, and further improve the practicality of the AI model. In addition, by sending the first content, the first location information, and the first task information to the server, the terminal device can execute corresponding processing tasks according to these information by the server and obtain corresponding task processing results, and the terminal device can obtain the task processing results from the server. Therefore, it is possible to reduce or avoid the workload of task processing on the electronic device side, and further improve the overall processing efficiency of the electronic device side.

[0051] In a possible design, the first content includes a first sub-content, the first location is a location before or after the first sub-content, and the first reference content includes the first sub-content.

[0052] Optionally, if the first position is a position before the first sub - content, in the content input to the first model, the first position information and the first task information are before the first sub - content; if the first position is a position after the first sub - content, in the content input to the first model, the first position information and the first task information are after the first sub - content.

[0053] In a possible design, the first content further includes a second sub - content, the first position is a position between the first sub - content and the second sub - content, and the first reference content further includes the second sub - content.

[0054] Optionally, in the content input to the first model, the first position information and the first task information are between the first sub - content and the second sub - content.

[0055] In a possible design, the first content includes a third sub - content, and the third sub - content is located at the first position; the second content specifically includes the content obtained by the first model performing the first processing task on the third sub - content according to the first reference content.

[0056] In a possible design, the first position information includes start position information and end position information; wherein, the start position information is used to indicate the starting point of the first position, and the end position information is used to indicate the ending point of the first position; the third sub - content is located between the starting point and the ending point in the first content.

[0057] In a possible design, in the content input to the first model, the third sub - content is between the start position information and the end position information; and / or, in the content input to the first model, the first task information is before the start position information, or between the start position information and the third sub - content, or between the third sub - content and the end position information, or after the end position information.

[0058] In a possible design, the terminal device is further configured to: before sending the first content, the first position information, and the first task information to the server, obtain the first content and the first indication information; wherein, the first indication information is used to indicate to perform the first processing task on the first position in the first content; and determine the first position information and the first task information according to the first content and the first indication information.

[0059] In a possible design, the terminal device is further configured to: after obtaining the first content and the first indication information, and before determining the first location information and the first task information according to the first content and the first indication information, display the first content and the first indication information, and receive a first operation; wherein, the first operation is used to indicate: according to the first content, convert the first indication information into the first location information and the first task information.

[0060] In a possible design, when the terminal device obtains the first content and the first indication information, it is specifically configured to: obtain and display the first content through a first application, and determine the first indication information according to a second operation acting on the first content received; wherein, the second operation is used to indicate the first indication information; or, obtain the first content and the first indication information indicated by a third operation received through the first application; or, receive the first content and the first indication information sent by a second application through the first application; wherein, the second application is used to generate the first content and the first indication information.

[0061] Optionally, the terminal device is further configured to: before receiving the first content and the first indication information sent by the second application through the first application, obtain and display the first content through the second application, and determine the first indication information according to a fourth operation acting on the first content received.

[0062] Optionally, when the terminal device sends the first content, the first location information and the first task information to the server, it is specifically configured to: send the first content, the first location information and the first task information to the server through the first application.

[0063] In a possible design, the terminal device is further configured to: before sending the first content, the first location information and the first task information to the server, receive a third operation; wherein, the third operation is used to indicate to perform a processing task according to the first content, the first location information and the first task information. Optionally, the terminal device is further configured to: before receiving the third operation, display the first content, the first location information and the first task information.

[0064] In a possible design, the terminal device is further configured to: send the second location information and the second task information to the server; wherein, the second location information is used to indicate the second location in the first content, and the second task information is used to indicate to perform a second processing task for the second location; the server is further configured to: when inputting the first content, the first location information, and the first task information into the first model, input the second location information and the second task information into the first model; the second content further includes: the content obtained by the first model performing the second processing task for the second location according to the second reference content; wherein, the second reference content includes part or all of the content in the first content except the second location.

[0065] In a possible design, the server is further configured to: before inputting the first content, the first location information, and the first task information into the first model to determine the second content, obtain training data, and perform model training on a set model according to the training data to obtain the first model; wherein, the training data includes a third content, third location information, third task information, and a fourth content; wherein, the third location information is used to indicate the third location in the third content, and the third task information is used to indicate to perform a third processing task for the third location; the fourth content includes: the content obtained by performing the third processing task for the third location according to part or all of the content in the third content except the third location.

[0066] In a possible design, when the server obtains the training data, it is specifically configured to: obtain the third content and the second indication information; wherein, the second indication information is used to indicate to perform the third processing task for the third location in the third content; determine the third location information and the third task information according to the second indication information; input the third content and the second indication information into the second model to determine a fifth content; or, input the content at the third location in the third content and the second indication information into the second model to determine the fifth content; or, when there is no content at the third location, input the second indication information into the second model to determine the fifth content; wherein, the fifth content includes the content obtained by the second model performing the third processing task according to the input content; determine the fourth content according to the fifth content.

[0067] Optionally, when determining the fourth content based on the fifth content, the server is specifically configured to: use the fifth content as the fourth content, or modify the fifth content in response to a received sixth operation to obtain the fourth content; where the sixth operation is used to indicate modifying the fifth content to the fourth content. Wherein, the second model may be a set generative AI model.

[0068] In a possible design, the training data further includes fourth position information and fourth task information; where the fourth position information is used to indicate a fourth position in the third content, and the fourth task information is used to indicate performing a fourth processing task on the fourth position; the fourth content further includes: content obtained by performing the fourth processing task on the fourth position according to some or all of the content in the third content other than the fourth position.

[0069] In a possible design, when the server obtains training data, it is specifically configured to: obtain the third content, second indication information, and third indication information; where the second indication information is used to indicate performing the third processing task on the third position in the third content; the third indication information is used to indicate performing the fourth processing task on the fourth position in the third content; input the third content and the second indication information into a second model to determine a fifth content; or input the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or input the second indication information into the second model to determine the fifth content; where the fifth content includes the content obtained by the second model performing the third processing task according to the input content; input the third content and the third indication information into the second model to determine a sixth content; or input the content at the fourth position in the third content and the third indication information into the second model to determine the sixth content; or input the third indication information into the second model to determine the sixth content; where the sixth content includes the content obtained by the second model performing the fourth processing task according to the input content; determine the fourth content based on the fifth content and the sixth content.

[0070] Optionally, when the server determines the fourth content based on the fifth content and the sixth content, it is specifically configured to: use the content obtained by combining the seventh content determined based on the fifth content and the eighth content determined based on the sixth content as the fourth content; where the seventh content is the fifth content or the content obtained by modifying the fifth content in response to a received operation; the eighth content is the sixth content or the content obtained by modifying the sixth content in response to a received operation.

[0071] In a possible design, when the server performs model training on a set model according to the training data to obtain the first model, it is specifically configured to: input the training data into the set model to determine a ninth content; wherein, the ninth content includes the content obtained by the set model performing the third processing task on the third position in the third content; and adjust the set model according to the third content and the ninth content to obtain the first model.

[0072] In a possible design, the ninth content further includes the content obtained by the set model performing the fourth processing task on the fourth position in the third content.

[0073] Optionally, the set model is a model with a set network structure or a set generative AI model.

[0074] In a possible design, the first content includes at least one of the following: text, image, audio, video, code.

[0075] In a possible design, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.

[0076] In a possible design, the first model belongs to a generative artificial intelligence (AI) model.

[0077] In a third aspect, the present application provides a model training method applied to an electronic device. The method includes: obtaining training data; wherein, the training data includes first content, first position information, first task information, and second content; wherein, the first position information is used to indicate a first position in the first content, and the first task information is used to indicate performing a first processing task on the first position; the second content includes: the content obtained by performing the first processing task on the first position according to part or all of the content in the first content except the first position; and performing model training on a set model according to the training data to obtain the first model. Optionally, the electronic device can be a server or a terminal device.

[0078] In this method, in the training data, some or all of the content in the first content except for the first position can be used as the context content of the first position, and the second content includes the content obtained by performing the first processing task on the first position according to the context content. Therefore, the model trained based on this training data can have the ability to perform corresponding processing tasks on the position by combining the context content of the position in the content. And performing task processing on a position by combining the context content of the position can obtain a more accurate processing result. Therefore, by using the above model training method, the accuracy of the result output when the trained first model performs tasks can be higher, and it is easier to meet the task processing requirements, thereby improving the practicality of the trained first model. Therefore, by using the above model training method, a model with better task processing effect can be trained.

[0079] In a possible design, the obtaining of the training data includes: obtaining the first content and the first indication information; wherein, the first indication information is used to indicate performing the first processing task on the first position in the first content; determining the first position information and the first task information according to the first indication information; inputting the first content and the first indication information into a second model to determine the third content; or, inputting the content at the first position in the first content and the first indication information into the second model to determine the third content; or, when there is no content at the first position, inputting the first indication information into the second model to determine the third content; wherein, the third content includes the content obtained by the second model performing the first processing task according to the input content; determining the second content according to the third content.

[0080] Optionally, the determining the second content according to the third content includes: using the third content as the second content, or, in response to a received first operation, modifying the third content to obtain the second content; wherein, the first operation is used to indicate modifying the third content to the second content. Wherein, the second model can be a set generative AI model.

[0081] The above method for obtaining training data can be applied to the scenario of training a first model for processing single tasks. Through the above method, the electronic device can use a set second model to realize the acquisition of part of the training data used as the task processing result. Among them, by modifying the content obtained by the second model processing, that is, the third content, according to the user operation, the accuracy of the task processing result data in the training data can be improved by means of manual correction. Further, the accuracy of model training using the training data can be improved, thereby obtaining a first model with better processing effect.

[0082] In a possible design, the training data further includes second location information and second task information; wherein, the second location information is used to indicate a second location in the first content, and the second task information is used to indicate to perform a second processing task on the second location; the second content further includes: content obtained by performing the second processing task on the second location according to part or all of the content in the first content other than the second location.

[0083] Based on the above method, an electronic device can implement the training of a first model for unified multi-task processing.

[0084] In a possible design, the obtaining of the training data includes: obtaining the first content, first indication information, and second indication information; wherein, the first indication information is used to indicate to perform the first processing task on the first location in the first content; the second indication information is used to indicate to perform the second processing task on the second location in the first content; inputting the first content and the first indication information into a second model to determine a third content; or, inputting the content at the first location in the first content and the first indication information into the second model to determine the third content; or, inputting the first indication information into the second model to determine the third content; wherein, the third content includes the content obtained by the second model performing the first processing task according to the input content; inputting the first content and the second indication information into the second model to determine a fourth content; or, inputting the content at the second location in the first content and the second indication information into the second model to determine the fourth content; or, inputting the second indication information into the second model to determine the fourth content; wherein, the fourth content includes the content obtained by the second model performing the second processing task according to the input content; determining the second content according to the third content and the fourth content.

[0085] Optionally, the determining the second content according to the third content and the fourth content includes: using the content obtained by combining the fifth content determined according to the third content and the sixth content determined according to the fourth content as the second content; wherein, the fifth content is the third content or the content obtained by modifying the third content in response to a received operation; the sixth content is the fourth content or the content obtained by modifying the fourth content in response to a received operation.

[0086] The above method for obtaining training data can be applied to the scenario of training a first model for processing multiple tasks. Through the above method, the electronic device can use the set second model to collect some of the training data that is used as the task processing result. Among them, by modifying the third content and the fourth content obtained by processing the second model according to the user operation, the accuracy of the task processing result data in the training data can be improved by means of manual correction. Further, the accuracy of model training using the training data can be improved, and thus a first model with better processing effect can be obtained.

[0087] In a possible design, training the set model according to the training data to obtain the first model includes: inputting the training data into the set model to determine a seventh content; where the seventh content includes the content obtained by the set model performing the first processing task on the first position in the first content; and adjusting the set model according to the first content and the seventh content to obtain the first model.

[0088] In the above method, the electronic device can adjust the set model according to the deviation between the task processing result with higher accuracy, i.e., the seventh content, and the task processing result obtained by the set model, i.e., the second content, so that the task processing result of the set model approaches the more accurate task processing result, and thus a first model with better processing effect is trained.

[0089] In a possible design, the seventh content further includes the content obtained by the set model performing the second processing task on the second position in the first content. Combining the above method, a first model with better processing effect for processing multiple tasks can be trained.

[0090] In a possible design, the first content includes at least one of the following: text, image, audio, video, code.

[0091] The first content in the above method can include at least one type of content. Therefore, based on the above method, a model that can better process single tasks or multiple tasks of various types of content can be trained.

[0092] In a possible design, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.

[0093] Among them, the semantic mask image can clearly and intuitively represent the positions of different types of content in the image, which helps to improve the accuracy of position recognition and task processing for the image.

[0094] In a possible design, the set model is a model of a set network structure or a set generative AI model, and the first model belongs to a generative artificial intelligence (AI) model. Among them, the generative AI model can perform generative tasks. By applying the generative AI model to the above method, the range of tasks that the trained model can execute can be further expanded, and the task execution ability of the trained model can be improved.

[0095] In a fourth aspect, the present application provides an electronic device, which includes a memory and one or more processors; wherein, the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the one or more processors, the electronic device is enabled to execute the method described in the first aspect or any possible design of the first aspect, or execute the method described in the third aspect or any possible design of the third aspect.

[0096] In a fifth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on an electronic device, the electronic device is enabled to execute the method described in the first aspect or any possible design of the first aspect, or execute the method described in the third aspect or any possible design of the third aspect.

[0097] In a sixth aspect, the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions run on an electronic device, the electronic device is enabled to execute the method described in the first aspect or any possible design of the first aspect, or execute the method described in the third aspect or any possible design of the third aspect.

[0098] In a seventh aspect, the present application provides a chip system, which includes a processor and a memory, and instructions are stored in the memory; when the instructions are executed by the processor, the method described in the first aspect or any possible design of the first aspect is implemented, or the method described in the third aspect or any possible design of the third aspect is implemented. The chip system can be composed of chips or can include chips and other discrete devices.

[0099] For the beneficial effects of the second to fifth aspects above, reference can be made to the corresponding beneficial effects in the first aspect, the second aspect, or the third aspect above, and details will not be repeated here. Description of the Drawings

[0100] Figure 1 It is a schematic diagram of a task processing method;

[0101] Figure 2Schematic diagram of the hardware architecture of an electronic device provided by an embodiment of the present application;

[0102] Figure 3 Schematic diagram of the software architecture of an electronic device provided by an embodiment of the present application;

[0103] Figure 4a Schematic diagram of the architecture of a content processing system provided by an embodiment of the present application;

[0104] Figure 4b Schematic diagram of the architecture of another content processing system provided by an embodiment of the present application;

[0105] Figure 5 Schematic diagram of a content processing method provided by an embodiment of the present application;

[0106] Figure 6 Schematic diagram of a method for editing content provided by an embodiment of the present application;

[0107] Figure 7 Schematic diagram of a content conversion method provided by an embodiment of the present application;

[0108] Figure 8 Schematic diagram of the result of content processing provided by an embodiment of the present application;

[0109] Figure 9 Schematic flow diagram of a content processing method provided by an embodiment of the present application;

[0110] Figure 10 Schematic diagram of the original content provided by an embodiment of the present application;

[0111] Figure 11 Schematic diagram of a content conversion method provided by an embodiment of the present application;

[0112] Figure 12 Schematic diagram of the content processing process provided by an embodiment of the present application;

[0113] Figure 13 Schematic diagram of a content conversion control method provided by an embodiment of the present application;

[0114] Figure 14 Schematic diagram of a content conversion control interface provided by an embodiment of the present application;

[0115] Figure 15 Schematic diagram of a content conversion control interface provided by an embodiment of the present application;

[0116] Figure 16 Schematic diagram of a content conversion control interface provided by an embodiment of the present application;

[0117] Figure 17 Schematic diagram of a content processing method provided by an embodiment of this application;

[0118] Figure 18 Schematic diagram of a content processing method provided by an embodiment of this application;

[0119] Figure 19 Schematic diagram of the structure of an electronic device provided by an embodiment of this application. Detailed implementation manners

[0120] In order to make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0121] Among them, in the description of the embodiments of this application, hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0122] For ease of understanding, illustrative descriptions of concepts related to this application are given for reference.

[0123] 1) An electronic device may be a device with processing and computing capabilities. For example, it may be a device with an AI model (such as a generative AI model).

[0124] In some embodiments of this application, the electronic device may be a computing device such as a server. For example, the electronic device may be a cloud server or the like.

[0125] In some embodiments of the present application, the electronic device may also be a portable device, such as a mobile phone, a tablet computer, a wearable device with wireless communication function (such as a watch, a bracelet, etc.), a vehicle-mounted terminal device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart home device (such as a smart TV, a smart speaker, etc.), a smart robot, a workshop device, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, or a wireless terminal in smart home, a flying device (such as a smart robot, a drone, an airplane), etc.

[0126] Among them, the wearable device is a portable device that users can directly wear on their bodies or integrate into their clothes or accessories.

[0127] In some embodiments of the present application, the electronic device may also be a portable terminal device that further includes other functions such as a personal digital assistant and / or a music player function. Exemplary embodiments of the portable terminal device include, but are not limited to, those equipped with or other operating systems. The above portable terminal device may also be other portable terminal devices, such as a laptop with a touch-sensitive surface (such as a touch panel), etc. It should also be understood that in some other embodiments of the present application, the above electronic device may not be a portable terminal device, but a desktop computer with a touch-sensitive surface (such as a touch panel).

[0128] 2) The generative artificial intelligence (AI) model is an AI model based on deep learning technology that can simulate human creative thinking and generate content such as text, images, audio, video, code, etc. with a certain degree of logic and coherence. The generative AI model can receive inputs such as text, images, audio, video, and code and generate new content in any of the above forms.

[0129] It should be understood that in the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one (item) of a, b, or c may represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c may be single or multiple.

[0130] Currently, the task execution effects of some AI models (such as generative AI models, large language foundation models, etc.) are still unstable. When performing certain tasks (such as complex tasks or associated tasks), it may not be possible to obtain good task processing results, resulting in low practicality of generative AI models. In addition, some AI models have good effects in processing a single task with a single input sequence input, but have poor effects in processing complex tasks with multiple tasks in a single input sequence input.

[0131] For example, as Figure 1 shown in the schematic diagram (a) in Figure 1 when the complex task set by the user for a piece of text includes multiple editing tasks annotated by the user for the piece of text (such as

[0132] the 5 tasks shown in the schematic diagram (a) in Figure 1 ), if the piece of text and the corresponding multiple editing tasks are directly input into the AI model as a single sequence, it is very difficult for the output result of the AI model to simultaneously meet all the requirements of the multiple editing tasks, so the processing effect is poor. Figure 1 Currently, a multi-task splitting scheme can also be adopted, that is, a complex task including multiple tasks is split into multiple single tasks that can be solved in a single input, and then the single tasks are respectively input into the AI model for processing, and then the processing results of each task are obtained respectively. For example, this scheme can be adopted to split the complex task shown in the schematic diagram (a) in

[0133] In the above method, after the complex task is split into individual tasks, the AI model can only process the tasks based on the limited content corresponding to the individual tasks, which may not guarantee the accuracy of processing and result in a poor processing effect. In some cases, there may be a relationship between multiple tasks (such as the "same as above" editing task commonly used in text annotation), and the split processing will also affect the processing effect. In addition, the above method requires each task to be processed separately, so the overall processing efficiency is low.

[0134] In summary, the current effect of using the AI model to process tasks is poor, especially the processing effect of using the AI model to process multiple tasks is poor, and the efficiency is low. Therefore, the practicality of the AI model is low.

[0135] Based on the above problems, in order to improve the practicality of the AI model, the embodiments of the present application provide a content processing method, system and electronic device. This solution can use the AI model to process tasks more simply, efficiently and accurately, thereby improving the efficiency and processing effect of the AI model in processing tasks and improving the practicality of the AI model.

[0136] The technical solution provided by the embodiments of the present application can be executed by any computing device with processing and computing capabilities, or can be executed by a system composed of multiple computing devices with processing, computing and communication capabilities. Among them, the computing device can be an electronic device, etc. The performance introduction of the electronic device can refer to the description in the above concept explanation or the relevant description in the following text.

[0137] In the following, the application of the technical solution of the present application in an electronic device or in a system including multiple electronic devices is used as an example for description. The implementation process in other computing devices is similar and will not be repeated. Optionally, the electronic device may have an AI model.

[0138] Next, refer to Figure 2 to introduce the structure of the electronic device to which the method provided by the embodiments of the present application is applicable.

[0139] As Figure 2 shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a SIM card interface 195, etc.

[0140] Among them, the sensor module 180 may include a gyroscope sensor, an acceleration sensor, a proximity light sensor, a fingerprint sensor, a touch sensor, a temperature sensor, a pressure sensor, a distance sensor, a magnetic sensor, an ambient light sensor, a barometric pressure sensor, a bone conduction sensor, etc.

[0141] It can be understood that Figure 2 the illustrated electronic device 100 is merely an example, which does not constitute a limitation on the electronic device, and the electronic device may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 2 The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.

[0142] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or may be integrated in one or more processors. Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0143] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data just used or recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0144] The execution of the content processing method provided by the embodiments of the present application can be controlled by the processor 110 or other components can be called to complete it. For example, the processing program of the embodiments of the present application stored in the internal memory 121 is called, or the processing program of the embodiments of the present application stored in a third-party device is called through the external memory interface 120 to control the wireless communication module 160 to perform data communication with other devices, improving the intelligence and convenience of the electronic device 100 and enhancing the user experience. The processor 110 can include different devices. For example, when the CPU and GPU are integrated, the CPU and GPU can cooperate to execute the content processing method provided by the embodiments of the present application. For example, some algorithms in the content processing method are executed by the CPU, and another part of the algorithms are executed by the GPU to obtain a faster processing efficiency.

[0145] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1. The display screen 194 can be used to display information input by the user or information provided to the user and various graphical user interfaces (GUIs). For example, the display screen 194 can display photos, videos, web pages, or files, etc.

[0146] In the embodiments of the present application, the display screen 194 can be an integrated flexible display screen, or a splicing display screen composed of two rigid screens and a flexible screen located between the two rigid screens.

[0147] The camera 193 (front camera or rear camera, or a single camera that can function as both a front camera and a rear camera) is used to capture still images or videos. Generally, the camera 193 may include a photosensitive element such as a lens group and an image sensor. Among them, the lens group includes multiple lenses (convex lenses or concave lenses) for collecting the optical signals reflected by the object to be photographed and transmitting the collected optical signals to the image sensor. The image sensor generates the original image of the object to be photographed based on the optical signals.

[0148] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, the code of application programs (such as the functions corresponding to the solution of the present application, etc.). The data storage area can store the data created during the use of the electronic device 100.

[0149] The internal memory 121 can also store one or more computer programs corresponding to the algorithms of the solution of the present application. The one or more computer programs are stored in the above internal memory 121 and are configured to be executed by one or more processors 110. The one or more computer programs include instructions, and the above instructions can be used to execute the respective steps in the following embodiments.

[0150] In addition, the internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0151] Of course, the code of the algorithm of the solution of the embodiment of the present application can also be stored in an external memory. In this case, the processor 110 can run the code of the algorithm of the solution of the present application stored in the external memory through the external memory interface 120.

[0152] The touch sensor, also known as the "touch panel". The touch sensor can be disposed on the display screen 194, and the touch sensor and the display screen 194 form a touch display screen, also known as the "touch screen". The touch sensor is used to detect touch operations acting on it or nearby. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor can also be disposed on the surface of the electronic device 100 at a different position from the display screen 194.

[0153] The wireless communication function of the electronic device 100 can be implemented by Antenna 1, Antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0154] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0155] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by Antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through Antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be disposed in the same device. In the embodiments of the present application, the mobile communication module 150 can also be used for information interaction with other devices.

[0156] The modulation and demodulation processor can include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to Speaker 170A, Receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor can be an independent device. In some other embodiments, the modulation and demodulation processor can be independent of the processor 110 and be disposed in the same device as the mobile communication module 150 or other functional modules.

[0157] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (WiFi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation. In the embodiments of the present application, the wireless communication module 160 may be used to establish connections with other electronic devices and perform data interaction. Or the wireless communication module 160 may be used to access an access point device, send control instructions to other electronic devices, or receive data sent from other electronic devices.

[0158] In addition, the electronic device 100 may implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc. The electronic device 100 may receive inputs from the keys 190 and generate key signal inputs related to the user settings and function controls of the electronic device 100. The electronic device 100 may use the motor 191 to generate vibration prompts (such as incoming call vibration prompts). The indicator 192 in the electronic device 100 may be an indicator light, which can be used to indicate the charging state, the change in battery level, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 in the electronic device 100 is used to connect the SIM card. The SIM card can be in contact with and separated from the electronic device 100 by being inserted into or removed from the SIM card interface 195.

[0159] It should be understood that in actual applications, the electronic device 100 may include more than Figure 2More or fewer components shown are not limited in the embodiments of the present application. The illustrated electronic device 100 is merely an example, and the electronic device 100 may have more or fewer components than those shown in the figure, two or more components may be combined, or different component configurations may be provided. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.

[0160] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. Communication between layers is through software interfaces. Exemplarily, as Figure 3 shown, the software architecture may be divided into four layers, from top to bottom are the application layer, the application framework layer (framework, FWK), the runtime and system libraries, and the (Linux) kernel layer.

[0161] The application layer is the top layer of the operating system, including the native applications of the operating system, such as the camera, gallery, calendar, Bluetooth, music, video, messages, etc., and may also include third-party applications. The application involved in the embodiments of the present application is abbreviated as application (APP), which is a software program capable of implementing one or more specific functions. Generally, multiple applications may be installed in the electronic device, such as a camera application, an email application, etc. The application mentioned hereinafter may be a system application pre-installed when the electronic device leaves the factory, or a third-party application downloaded from the network or obtained from other electronic devices during the use of the electronic device by the user.

[0162] Certainly, for developers, developers may write application programs and install them in this layer. In one possible implementation, the application program may be developed using the Java language and completed by calling the application programming interface (API) provided by the application framework layer. Developers may interact with the underlying layer of the operating system (such as the kernel layer, etc.) through the application framework to develop their own application programs.

[0163] The application framework layer is the API and programming framework for the application layer. The application framework layer may include some predefined functions. The application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.

[0164] The window manager is used to manage window programs. The window manager may obtain the display screen size, determine whether there is a status bar, lock the display screen (or screen), capture the display screen, etc.

[0165] The content provider is used to store and retrieve data and make this data accessible to applications. The data may include information such as files (e.g., documents, videos, images, audio), text, etc.

[0166] The view system includes visual controls, such as controls for displaying content like text, pictures, documents, etc. The view system can be used to build applications. The interface in the display window can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.

[0167] The telephony manager is used to provide the communication functions of the electronic device. The notification manager enables applications to display notification information in the status bar, can be used to convey informative messages, and can automatically disappear after a short stay without user interaction.

[0168] The runtime includes a core library and a virtual machine. The runtime is responsible for the scheduling and management of the system.

[0169] The core library of the system consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of the system. The application layer and the application framework layer run in the virtual machine. Taking Java as an example, the virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0170] The system library can include multiple functional modules. For example: the surface manager, the media library, the 3D graphics processing library (e.g., OpenGL ES), the 2D graphics engine (e.g., SGL), the image processing library, etc. The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.564, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is the drawing engine for 2D drawing.

[0171] The kernel layer provides the core system services of the operating system. Services such as security, memory management, process management, network protocol stack, and driver model are all implemented based on the kernel layer. The kernel layer also serves as an abstraction layer between the hardware and the software stack. There are many driver programs related to electronic devices in this layer. The main drivers include: display driver; keyboard driver as an input device; Flash driver based on memory technology devices; camera driver; audio driver; Bluetooth driver; WiFi driver, etc.

[0172] It should be understood that the functional services described above are only examples. In actual applications, electronic devices can also be divided into more or fewer functional services according to other factors, or the functions of each service can be divided in other ways, or the functional services can be not divided at all but work as a whole.

[0173] The solution provided by the embodiments of the present application will be described in detail below.

[0174] Figure 4a It is a schematic diagram of the architecture of a content processing system provided by an embodiment of the present application. This content processing system can be deployed in an electronic device. As Figure 4a shown, the content processing system may include: a content acquisition unit, a content editing unit, a content conversion unit, an AI model, and a display unit.

[0175] Among them, the electronic device can be a server or a terminal device.

[0176] Exemplarily, the server can be a cloud server.

[0177] Exemplarily, when the electronic device is a terminal device, the hardware architecture of the terminal device can be implemented using the architecture shown in Figure 2 and the software architecture of the terminal device can be implemented using the architecture shown in Figure 3 .

[0178] Exemplarily, when the software architecture of the terminal device is implemented using the architecture shown in Figure 3 , the functional units in this content processing system can be respectively deployed in the architecture layers such as the application layer or the application framework layer shown in Figure 3 , and there is no specific limitation in the embodiments of the present application.

[0179] The content acquisition unit can be used to acquire the original content to be processed. The format of the original content to be processed can be various media formats such as text, image, audio, video, code, etc. The content acquisition unit can also be used to send the original content to the content editing unit. Among them, the original content can be understood as the content before editing.

[0180] The content editing unit can be used to interact with the user, so as to perform task editing on the original content according to the user's operations. Among them, the user can perform editing operations on any one or any multiple positions of the original content to indicate tasks to be processed for any one or any multiple positions of the original content. Among them, there may or may not be content at each position. The absence of content can also be understood as the content being empty. The content editing unit can add task indication information for indicating the task corresponding to the editing operation to each position according to each editing operation performed by the user on each position of the original content, and then obtain the edited content. In some embodiments of the present application, the edited content may include: the original content, and the task indication information for each of one or more tasks. Among them, each of the one or more tasks may specifically be: a task (or operation) performed on one position in the original content, and the task indication information for each task is used to indicate the task. In the above method, the original content may include one or more sub - contents, and each sub - content may be part or all of the original content. Each position in the original content may include a sub - content, or may not include a sub - content (which can also be understood as the sub - content included at that position being empty). Different sub - contents in the original content may be completely the same, partially the same, or completely different. Among them, when different sub - contents in the original content are completely the same, the tasks corresponding to different sub - contents are different; when different sub - contents in the original content are partially the same or completely different, the tasks corresponding to different sub - contents may be the same or different. The content editing unit can also be used to send the edited content to the content conversion unit.

[0181] In some embodiments of the present application, the above one or more tasks may be tasks indicated by one user, or tasks indicated by multiple users.

[0182] The content conversion unit can be used to convert the edited content into single - task prompt content or multi - task prompt content. Among them, when the edited content includes one task, the content conversion unit can convert the edited content into single - task prompt content; when the edited content includes multiple tasks, the content conversion unit can convert the edited content into multi - task prompt content. In some embodiments of the present application, the single - task prompt content may include: the original content, and the task prompt content corresponding to the above one task. The multi - task prompt content may include: the original content, and the task prompt content corresponding to each of the above multiple tasks. Among them, the task prompt content corresponding to each task may include: the sub - content at the position corresponding to the task, the position information of the position corresponding to the task, and the task indication information of the task. Among them, the position information is used to indicate the position. The content conversion unit can also be used to send the single - task prompt content or multi - task prompt content to the AI model.

[0183] The AI model can be used to generate single-task processing results based on single-task prompt content, or generate multi-task processing results based on multi-task prompt content. Among them, the single-task processing result includes the content obtained after performing one of the above tasks on the original content. The multi-task processing result includes the content obtained after performing multiple of the above tasks on the original content. The single-task prompt content or multi-task prompt content can be the input content of the AI model, and the single-task processing result or multi-task processing result can be the output content of the AI model. In some embodiments of the present application, the AI model can be a trained model capable of recognizing and processing single-task prompt content or multi-task prompt content. Among them, regarding the training process of the AI model, reference can be made to the relevant descriptions in the following embodiments, which will not be elaborated here for the time being. In some other embodiments of the present application, when the AI model is a model that cannot directly recognize and process multi-task prompt content, the content conversion unit can convert the multi-task prompt content into multiple single-task prompt content, and then input each single-task prompt content into the AI model for processing respectively, and obtain the single-task processing results corresponding to each single-task prompt content output by the AI model, and then the content obtained after performing multiple of the above tasks on the original content can be summarized.

[0184] The display unit can be used to display the original content, or the edited content, or the single-task prompt content, or the multi-task prompt content, or the single-task processing result, or the multi-task processing result. In some embodiments of the present application, the display unit can be used to display the original content after the content acquisition unit acquires the original content. Or, it can be used to display the edited content after the content editing unit obtains the edited content. Or, it can be used to display the single-task prompt content or multi-task prompt content after the content conversion unit converts the edited content into single-task prompt content or multi-task prompt content. Or, it can be used to display the single-task processing result or multi-task processing result after the AI model outputs the single-task processing result or multi-task processing result.

[0185] Figure 4b It is a schematic diagram of the architecture of another content processing system provided by the embodiments of the present application. As Figure 4b shown in the figure, the content processing system may include a terminal device and a server. Optionally, the terminal device may include a content acquisition unit, a content editing unit, a content conversion unit, a display unit, and a communication unit. The server may include a communication unit and an AI model.

[0186] Among them, the content acquisition unit, content editing unit, content conversion unit, and display unit in the terminal device can be referred to the above description and will not be repeated here. The communication unit in the terminal device can be used to send the single-task prompt content or multi-task prompt content determined by the content conversion unit to the communication unit in the server, and can also be used to send the single-task processing result or multi-task processing result from the communication unit in the server to other units. The communication unit in the server can be used to send the received single-task prompt content or multi-task prompt content to the AI model so that the AI model processes the single-task prompt content or multi-task prompt content. The communication unit in the server can also send the single-task processing result or multi-task processing result determined by the AI model to the communication module of the terminal device. Regarding the AI model in the server, it can be referred to the above description and will not be repeated here.

[0187] The above Figure 4b The content processing system shown in Figure 4a is different from the content processing system shown in the above Figure 4a in that the content processing system in Figure 4b is deployed on the same electronic device, while the content processing system shown in Figure 4b is distributedly deployed on different electronic devices. In addition, Figure 4a the architecture and the structure and functions of the internal functional units of the content processing system shown in can be implemented with reference to the content processing system shown in

[0188] It should be understood that the architecture of the content processing system described above is only an example. In practical applications, the electronic device can also be divided into more or fewer functional units according to other factors, or the functions of each unit can be divided in other ways, or the functional units can not be divided and the whole can work.

[0189] Based on the above description, taking the content processing method provided in the embodiments of the present application as an example applied to the same electronic device and taking the scenario where the electronic device uses the AI model to process multi-tasks as an example, a possible content processing method provided in the embodiments of the present application can be referred to Figure 5 . Regarding the scenario where the electronic device uses the AI model to process single tasks, it can be implemented with reference to the method shown in Figure 5 and will not be elaborated in the embodiments of the present application.

[0190] As Figure 5 shown, a possible content processing method may include:

[0191] S501: The electronic device acquires the original content and displays the original content.

[0192] In some embodiments of the present application, the original content may be content in various media formats such as documents, texts, images, audios, videos, codes, etc.

[0193] The manner in which the electronic device obtains the original content may be, for example, obtaining the original content input by the user, generating the original content according to the user's operation, obtaining the original content sent by other devices, etc. In the embodiments of the present application, the manner in which the electronic device obtains the original content is not specifically limited.

[0194] Example 1. In one example, taking the original content as content in text format as an example, the original content may be Figure 6 a section of text shown in the schematic diagram (a) in. Among them, *, x, X, etc. are all used to represent the text content.

[0195] Optionally, when the method provided in the embodiments of the present application is applied to Figure 4a the content processing system shown in, the step of obtaining the original content described in step S501 may be executed by Figure 4a the content acquisition unit shown in, and the step of displaying the original content described in step S501 may be executed by Figure 4a the display unit shown in.

[0196] S502: The electronic device generates the edited content according to multiple editing operations performed by the user on the original content, and displays the edited content.

[0197] In some embodiments of the present application, each editing operation performed by the user is used to select any position of the original content and indicate a task to be performed for that position. The edited content may include the original content and the task indication information corresponding to each editing operation in the multiple editing operations. Among them, the task indication information corresponding to each editing operation is used to indicate the task indicated by the editing operation, that is, the processing task to be performed for the position selected by the editing operation.

[0198] Optionally, the edited content may further include the content markers corresponding to each editing operation in the multiple editing operations. Among them, the content marker corresponding to each editing operation is used to mark the position selected by the editing operation.

[0199] Among them, when there is content at a position selected by the user in the original content, the user's selection of that position is equivalent to selecting the sub-content at that position. Therefore, in this scenario, selecting a position is equivalent to selecting the sub-content. Therefore, in the following embodiments of the present application, selecting the sub-content is equivalent to selecting the position where the sub-content is located.

[0200] In some embodiments of the present application, the positions or sub - contents selected by different editing operations can be exactly the same (which can also be understood as completely overlapping or coinciding), or can be partially the same (which can also be understood as partially overlapping or coinciding), or can be completely different (which can also be understood as completely non - overlapping or non - coinciding).

[0201] Based on the above - mentioned method, the original content can include multiple sub - contents.

[0202] In specific implementation, one editing operation performed by the user on the original content can include: selecting any position of the original content and performing an editing operation on this position, and this editing operation can be used to indicate a task to be executed for this position. Optionally, the sub - content at this position can be a subset of the original content. Among them, the user performing one editing operation is equivalent to issuing one task.

[0203] In some embodiments of the present application, the original content, position information, and task indication information in the edited content can exist independently; or some of the contents can be integrated into a whole, and the other part exists independently; or all of them can be integrated into a whole, and no specific limitation is made in the embodiments of the present application.

[0204] As an alternative implementation manner, every time the user performs an editing operation, the electronic device can, in response to this editing operation, add task indication information corresponding to this editing operation (which can be used to indicate the task corresponding to this editing operation) to the original content, and can display the content obtained after adding the task indication information. The user can perform multiple editing operations on the original content according to the above - mentioned method. After the multiple editing operations performed by the user are completed, the electronic device can obtain the edited content and can display the edited content.

[0205] Example 2: In an example, based on Example 1 above, when the user performs an editing operation on the original content shown in the schematic diagram (a) in Figure 6 as shown in the schematic diagram (b) in Figure 6 one editing operation performed by the user on the original content can include: selecting segment 1 in the original content and adding editing opinion 1 of "A processing" for segment 1 to indicate performing A processing on segment 1. Then the task corresponding to this editing operation is: performing A processing on segment 1. Among them, the sub - content corresponding to this editing operation is segment 1, and the task indication information can be the editing opinion 1 shown in the schematic diagram (b) in Figure 6 After the user performs the above - mentioned one editing operation on the original content shown in the schematic diagram (a) in Figure 6 the electronic device can switch from displaying the content shown in the schematic diagram (a) in Figure 6 to displaying the content shown in the schematic diagram (b) in Figure 6the content shown in the schematic diagram (b) therein, so as to display the editing operations performed by the user on the original content to the user. Based on the same method as described above, the user can perform multiple editing operations on the original content in sequence, and the electronic device can obtain and display the edited content after the user performs multiple editing operations.

[0206] In one example, the edited content displayed by the electronic device may be Figure 7 the content shown in the schematic diagram (a) therein. According to this content, it can be determined that the user has performed three editing operations. Among them, the sub-content corresponding to the first editing operation is fragment 1, the corresponding task indication information is editing opinion 1, and the corresponding task is: perform processing A on fragment 1. The sub-content corresponding to the second editing operation is fragment 2, the corresponding task indication information is editing opinion 2, and the corresponding task is: perform processing B on fragment 2. The sub-content corresponding to the third editing operation is fragment 3, the corresponding task indication information is editing opinion 3, and the corresponding task is: perform processing C on fragment 3.

[0207] Optionally, when the method provided in the embodiments of the present application is applied to Figure 4a the content processing system shown therein, the step of generating the edited content according to the multiple editing operations performed by the user on the original content in step S502 may be performed by Figure 4a the content editing unit shown therein, and the step of displaying the edited content in step S502 may be performed by Figure 4a the display unit shown therein.

[0208] S503: In response to the received conversion operation, the electronic device converts the edited content into multitask prompt content and displays the multitask prompt content.

[0209] Among them, the conversion operation is used to indicate converting the edited content into multitask prompt content. In some embodiments of the present application, the process of converting the edited content into multitask prompt content may also be referred to as creation.

[0210] In some embodiments of the present application, the multitask prompt content may include: the original content, the task indication information corresponding to each editing operation in the multiple editing operations on the original content, and the position information of the position corresponding to each editing operation in the multiple editing operations. Among them, the position corresponding to each editing operation is the position selected by the editing operation, and the position information of the position corresponding to each editing operation is used to indicate this position.

[0211] In other embodiments of the present application, the multitask prompt content may be the content obtained by adding the position information of each position to the edited content.

[0212] In the above method, any position, the task indication information corresponding to this position, and the position information of this position can form the task prompt content of this position. Therefore, the multi-task prompt content can include the task prompt content of each position among multiple positions.

[0213] In a possible solution, the position information of each position can include start position information and end position information. Among them, the start position information is used to indicate the starting point of this position, and the end position information can be used to indicate the end point (i.e., the ending point) of this position.

[0214] Optionally, the start position information and the end position information can be the same or different.

[0215] Optionally, the start position information and the end position information can be displayed as information in various formats such as marks, icons, texts, symbols, etc. In the embodiments of the present application, the formats of the start position information and the end position information are not specifically limited.

[0216] Optionally, in the multi-task prompt content, the arrangement order among any position, the start position information of this position, the end position information of this position, and the task indication information corresponding to this position can be arbitrary or can be preset. For example, in a possible way, when there is sub-content at this position, the sub-content can be located between the start position information and the end position information of this position. The task indication information corresponding to this position can be located between the start position information of this position and this position, or can be located between this position and the end position information of this position.

[0217] In another possible solution, the position information of each position can include marking information for marking this position. The marking information can be, for example, an underline, a contour line, a shadow, an icon, etc. In the embodiments of the present application, the format of the marking information is not specifically limited.

[0218] Example 3. In an example, based on Example 2 above, the multi-task prompt content obtained after the electronic device converts the edited content shown in the schematic diagram (a) in Figure 7 can be the content shown in the schematic diagram (b) in Figure 7 Specifically, in the edited content shown in the schematic diagram (a) in Figure 7 , the record of each editing operation of the user includes the corresponding segment and editing opinion. And in Figure 7In the multitask prompt content shown in the schematic diagram (b) therein, the record of each user editing operation includes the corresponding segment, the start position information and end position information of the segment, and the editing opinion. Among them, taking the sorting and format of the segment, the start position information and end position information of the segment, and the editing opinion in the multitask prompt content as "[start position information of the segment][editing opinion of the segment (i.e., task indication information of the segment)][segment][end position information of the segment]" as an example, such as Figure 7 As shown in the schematic diagram (b) therein, the task prompt content of segment 1 includes: [start position 1][processing A]xxxxxxxx[end position 1]. Among them, [start position 1] is the start position information of segment 1, [processing A] is the task indication information of segment 1, xxxxxxxx is the content included in segment 1, and [end position 1] is the end position information of segment 1. Since segment 2 and segment 3 overlap and segment 2 is before segment 3, the task prompt content of segment 2 includes the start position information and task indication information in the task prompt content of segment 3, and the task prompt information of segment 3 includes the end position information in the task prompt content of segment 2. Then, as Figure 7 As shown in the schematic diagram (b) therein, the combined display of the task prompt content of segment 2 and segment 3 is: [start position 2][processing B]xxxxxxx[start position 3][processing C]XXXX[end position 2]*******[end position 3]. Among them, [start position 2] is the start position information of segment 2, [processing B] is the task indication information of segment 2, xxxxxxx and XXXX are the content included in segment 2, [end position 2] is the end position information of segment 2, and these items form the task prompt content of segment 2. [start position 3] is the start position information of segment 3, [processing C] is the task indication information of segment 3, XXXX and ******* are the content included in segment 3, and [end position 3] is the end position information of segment 3, and these items form the task prompt content of segment 3.

[0219] In some embodiments of the present application, the position information of each position may be associated with the sub-content at that position. Alternatively, the position information of each position may include the identification information of that position. For example, in the Figure 7 multitask prompt content shown in the schematic diagram (b) therein, the serial number "1" in the start position information and end position information of segment 1 may be used as the identification information of segment 1, and the serial number "2" in the start position information and end position information of segment 2 may be used as the identification information of segment 2.

[0220] In some embodiments of the present application, when the position selected by the user during the editing of the original content does not include the valid content in the original content, it can be considered that the sub-content at this position is empty, and the task prompt content corresponding to this position may not include the sub-content either. That is, the task prompt content corresponding to the position may only include the editing opinion and position information corresponding to the position. Or, it can be considered that the sub-content at this position is the original content, and the task prompt content corresponding to the position includes the editing opinion corresponding to the position, the original content, and the position information of the original content. Or, it can be considered that the sub-content at this position is the content within the set range where the position for the user to edit is located, and the task prompt content corresponding to the position includes the editing opinion corresponding to the position, the content within the set range, and the position information of the content within the set range.

[0221] Optionally, when the method provided by the embodiments of the present application is applied to Figure 4a the content processing system shown in, the step of converting the edited content into multi-task prompt content in step S503 in response to the received first operation may be performed by Figure 4a the content conversion unit shown in, and the step of displaying the multi-task prompt content in step S503 may be performed by Figure 4a the display unit shown in.

[0222] S504: The electronic device generates a multi-task processing result according to the AI model and the multi-task prompt content.

[0223] Optionally, the AI model may be a generative AI model.

[0224] In some embodiments of the present application, the above AI model may be a trained AI model capable of recognizing multi-task prompt content. The training process of the AI model may refer to the method described in the following embodiments, which will not be elaborated here for the time being. The electronic device may directly input the multi-task prompt content as a single input sequence into the trained AI model and obtain the multi-task processing result output by the AI model.

[0225] As a possible solution, when the AI model processes the multi-task prompt content, it can determine the corresponding multiple tasks according to the multiple task prompt contents in the multi-task prompt content, and process each task separately according to the task prompt content of each task in the multiple tasks. When processing each task, the AI model can determine the task to be executed based on the editing opinion of the task, and can determine the position corresponding to the task and the context content of that position (i.e., the content in the original content other than the selected position) based on the position information corresponding to the task, and then combine the context content of that position to execute the task to be executed for that position. Based on this method, the AI model can also refer to more content when performing single-task processing, thereby improving the processing accuracy and the processing effect of the AI model.

[0226] Example 4. In one example, based on the above Example 3, taking the multi-task prompt content shown in the schematic diagram (b) in Figure 7 as an example, after the electronic device inputs the multi-task prompt content shown in the schematic diagram (b) in Figure 7 into the AI model, the multi-task processing result output by the AI model can be the content shown in Figure 8 , and this content includes the content obtained after performing corresponding tasks on segments 1 to 3 shown in the schematic diagram (b) in Figure 7 respectively. Among them, Figure 8 the segment 10 shown in is the segment obtained after performing the task of "performing A processing on segment 1" on segment 1 shown in the schematic diagram (b) in Figure 7 . Figure 8 The segment 20 shown in is the segment obtained after performing the task of "performing B processing on segment 2" on segment 2 shown in the schematic diagram (b) in Figure 7 , Figure 8 and the segment 30 shown in is the segment obtained after performing the task of "performing C processing on segment 3" on segment 3 shown in the schematic diagram (b) in Figure 7 .

[0227] Optionally, when the method provided in the embodiments of the present application is applied to the content processing system shown in Figure 4a , step S504 can be executed by the AI model shown in Figure 4a .

[0228] In the above method, the electronic device can convert the content containing multiple tasks to be executed (i.e., the edited content) into multi-task prompt content based on a certain format, and can describe and prompt the multiple tasks to be executed respectively through the relevant content in the multi-task prompt content, so as to facilitate the recognition of the AI model. The electronic device can directly input the multi-task prompt content containing multiple tasks into the AI model as a single input sequence for processing, and obtain the corresponding multi-task processing results. It can achieve the effect of executing multiple tasks in a single input sequence while ensuring the processing effect, that is, obtaining the output results of multiple tasks through one input, thereby improving the processing efficiency.

[0229] Next, the method described in the above embodiments will be exemplarily described with specific examples.

[0230] Taking the original content described in the above embodiments as the text content in the original document as an example, as Figure 9 shown in, the process of a content processing method provided by an embodiment of the present application may include:

[0231] Step 1: The electronic device obtains the original document.

[0232] Among them, the original document includes the original content.

[0233] Example 5: In one example, the original content included in the original document may be Figure 10 the text content shown in.

[0234] Step 2: The electronic device determines the edited document according to the interactive editing performed by the user on the original document.

[0235] Among them, the interactive editing may include the selection of multiple segments in the original document and the generation of corresponding editing opinions for each selected segment. The editing opinions are used to indicate the tasks to be executed for the segments.

[0236] The edited document obtained by the user's interactive editing of the original document includes the edited content.

[0237] Among them, a segment may be a subset of the text content in the original document, and there may be intersections among multiple segments. Each segment may correspond to at least one editing opinion. Each editing opinion corresponds to a task.

[0238] Exemplarily, as Figure 9 shown in, the multiple segments may include segment 1 to segment n (n is a positive integer greater than 1), and the editing opinions corresponding to segment 1 to segment n are editing opinion 1 to editing opinion n respectively.

[0239] Example 6: In one example, based on the above Example 5, the user Figure 10The content obtained by interactively editing the text content shown in is the edited content, such as Figure 11 as shown in the schematic diagram (a) in. Among them, the edited content includes 5 editing opinions for 5 segments in the original content.

[0240] Step 3: The electronic device generates multitask prompt text according to the edited content.

[0241] Among them, the multitask prompt text can be the content obtained by marking the start and end positions of the edited segments and the corresponding editing opinions on the original content in the original document.

[0242] Each edited segment in the multitask prompt text has a corresponding start symbol, editing opinion, and end symbol. Among them, the start symbol is used as the start position information of the corresponding segment, and the end symbol is used as the termination position information of the corresponding segment. Each segment and the corresponding start symbol, editing opinion, and end symbol form the task prompt content corresponding to the segment.

[0243] Exemplarily, as Figure 9 shown in, the start symbol of the edited segment 1 can be start symbol 1, the editing opinion of segment 1 can be editing opinion 1, and the end symbol of segment 1 can be end symbol 1. The start symbol of the edited segment 2 can be start symbol 2, the editing opinion of segment 2 can be editing opinion 2, and the end symbol of segment 2 can be end symbol 2. And so on for other segments.

[0244] Example 7. In one example, based on Example 6 above, Figure 11 the multitask prompt text obtained by converting the edited content shown in the schematic diagram (a) in can be Figure 11 the multitask prompt text shown in the schematic diagram (b) in. Among them, Figure 11 in the edited content shown in the schematic diagram (a) in, each editing record of the user includes a segment area and an editing opinion. The electronic device can convert each editing record of the user into "[ / Prmpt <id><Editor's Comment>]Original Fragment[ / Prmpt <id>in the format of ” to obtain Figure 11 the multi-task prompt text shown in the schematic diagram (b) in [ / Prmpt <id><Edit Comment>]Indicates the starting position of the document segment to be edited and the corresponding edit comment at the beginning,[ / Prmpt <id>As an end, it indicates the end position of the document fragment to be edited. The original fragment in the middle includes the selected fragment to be edited by the user, and this fragment is a subset of the original content in the original document. For example, in Figure 11 in the multitask prompt text shown in the schematic diagram (b), one of the edited fragments can include "Just as Confucius said: 'Consider others in the light of your own.'" The task prompt content corresponding to this fragment can be Figure 11 in the schematic diagram (b) shown in "[ / Prmpt<1><The allusion is used incorrectly, please modify>]Just as Confucius said: 'Consider others in the light of your own.'[ / Prmpt<1>]", where the first symbol " / Prmpt<1>" in the task prompt content is the start symbol and can be used as the start position information of the edited fragment, and the second symbol " / Prmpt<1>" is the end symbol and can be used as the termination position information of the edited fragment. And so on for other fragments.

[0245] Step 4: The electronic device inputs the multitask prompt text into the AI model to obtain a multitask processing result.

[0246] In a possible solution, for an AI model that can directly recognize the multitask prompt text (such as the trained AI model described in the foregoing embodiments), the electronic device does not need to perform other processing and directly inputs the obtained multitask prompt text into the AI model to obtain the multitask processing result output by the AI model.

[0247] In another possible solution, for an AI model that cannot recognize the multitask prompt text, the electronic device can first generate multiple single-task prompt texts according to the multitask prompt text, and respectively input each single-task prompt text into the AI model for processing to obtain the task processing result corresponding to each single-task prompt text. Then, by summarizing and integrating the task processing results corresponding to the multiple single-task prompt texts, the multitask processing result corresponding to the multitask prompt text can be obtained. Wherein, each single-task prompt text can include the task prompt content of a fragment corresponding to an editing opinion. Optionally, the fragment in the task prompt content (i.e., the fragment corresponding to the editing opinion) can also be replaced with the content of the paragraph where the fragment is located or the original content.

[0248] In yet another possible solution, for an AI model that cannot recognize the multitask prompt text, the electronic device can prompt in the system setting or user context to explain how to process the multitask prompt text, thereby assisting the AI model to better process the multitask prompt text.

[0249] Exemplarily, as Figure 9 shown, the multitask processing result output by the AI model after processing the multitask prompt text can include the content obtained after performing corresponding processing on each fragment. For example, Figure 9 The new segments 1 to n shown therein. Among them, the new segments 1 to n are the contents obtained by the AI model after performing corresponding task processing on segments 1 to n respectively.

[0250] The above method realizes more efficient document processing. Among them, the electronic device makes the AI model execute tasks corresponding to multiple editing opinions in one input (i.e., a single input sequence) by converting the document, so the processing efficiency is relatively high. In addition, compared with the method of decomposing complex tasks into multiple executable simple tasks, the above method has the characteristic of richer context, so the processing effect is better.

[0251] Taking the original content described in the above embodiment as the original image as an example, as Figure 12 shown therein, the process of a content processing method provided by an embodiment of the present application may include:

[0252] Step 1: The electronic device acquires the original image.

[0253] Exemplarily, the original image may be Figure 12 the original image shown therein.

[0254] Step 2: The electronic device determines the edited image according to the interactive editing performed by the user on the original image.

[0255] Among them, the interactive editing may include the selection of multiple image regions in the original image and the editing opinions generated for each selected image region.

[0256] The edited image obtained after the user performs interactive editing on the original image includes the edited content.

[0257] Among them, the image region may be a subset of the original image, and there may be overlaps between multiple image regions. Each image region may correspond to at least one editing opinion. Each editing opinion corresponds to a task.

[0258] Exemplarily, the user performs interactive editing on Figure 12 the original image shown therein, and the obtained image may be Figure 12 the edited image shown therein. Exemplarily, the regions selected by the user in the edited image may include a background region and a portrait region. Among them, the editing opinion 1 corresponding to the background region may be: change the background. The editing opinion 2 corresponding to the portrait region may be: change the hairstyle.

[0259] Step 3: The electronic device generates multi-task prompt content according to the edited image.

[0260] Exemplarily, as Figure 12 As shown in [figure reference], the multitask prompt content generated by the electronic device based on the edited image may include: the original image, the semantic mask image obtained by performing semantic segmentation on the original image, and the editing opinions, namely editing opinion 1 and editing opinion 2. Among them, the semantic mask image includes mask area 1 and mask area 2. Mask area 1 is the position information of the background area, which can be used to indicate the position of the background area in the original image. Mask area 2 is the position information of the portrait area, which can be used to indicate the position of the portrait area in the original image. Editing opinion 1 is the editing opinion corresponding to the background area, and editing opinion 2 is the editing opinion corresponding to the portrait area.

[0261] Optionally, the above-mentioned editing opinions can be separate text contents or marker contents added in the mask areas of the semantic mask image. In the embodiments of the present application, the format and setting method of the editing opinions are not specifically limited.

[0262] Step 4: The electronic device inputs the multitask prompt content into the AI model to obtain the multitask processing result.

[0263] Exemplarily, the content obtained after processing the multitask prompt content by the AI model can be Figure 12 the multitask processing result image shown in [figure reference]. Among them, the multitask processing result image contains the content obtained by processing the background area and the portrait area in the original image according to the corresponding editing opinions respectively.

[0264] In the above method, for the image content, the user can select different image areas during the editing process and specify different editing tasks for each area. The electronic device converts the image, enabling the AI model to execute the processing of multiple editing opinions for the image in one input and generate the processing results of multiple image editing tasks, so the processing efficiency is relatively high.

[0265] Based on the above description, the embodiments of the present application further provide a content conversion control method, which can be applied to the foregoing embodiments, specifically in the process of the electronic device executing steps S501 - S503 in the foregoing embodiments.

[0266] Referring to Figure 13 , the content conversion control method provided by the embodiments of the present application can be any of the following methods:

[0267] Method 1: The electronic device executes the above steps S501 - S503 through the first application.

[0268] In this method, the first application can generate or obtain the original content and display the original content. After displaying the original content, the first application can generate the edited content in response to the user's operation of editing the original content and display the edited content. After displaying the edited content, the first application can generate multitasking prompt content based on the edited content in response to the received conversion operation and display the multitasking prompt content.

[0269] In this method, the first application can be installed in the electronic device. Optionally, the first application can be an application for generating the original content.

[0270] In some embodiments of the present application, when the first application displays the edited content, it can display a conversion control for triggering content conversion. And it can generate multitasking prompt content based on the edited content and display the multitasking prompt content in response to the received conversion operation acting on the conversion control. Among them, the conversion operation is used to indicate converting the edited content into multitasking prompt content.

[0271] Exemplarily, taking the original content as the text content in the original document in the above embodiment, the first application can be a document application. As Figure 14 shown, when the document application displays the edited content, it can display a conversion control and, in response to the user's operation of clicking the conversion control, convert the edited content into multitasking prompt content and display the multitasking prompt content.

[0272] Optionally, the way for the first application to display the multitasking prompt content can be: displaying the multitasking prompt content after the original content, or replacing the original content with the multitasking prompt content, or displaying the multitasking prompt content in a new interface or a newly created document interface.

[0273] Method 2: The electronic device executes steps S501 - S502 through the first application and executes step S503 through the second application.

[0274] Among them, the second application can be a system application or a third - party application. Exemplarily, the system application can be a system - type AI processing application such as a smart assistant. The third - party application can be a third - party AI processing application installed in the electronic device, etc.

[0275] In this method, the first application can generate or obtain the original content and display the original content. After the first application displays the original content, it can generate the edited content in response to the user's operation of editing the original content and display the edited content. The second application can convert the edited content into multitasking prompt content in response to the received conversion operation after obtaining the edited content.

[0276] Among them, the way for the second application to obtain the edited content can be any of the following:

[0277] 1) After the first application displays the edited content, it can forward the edited content to the second application in response to the received forwarding operation.

[0278] For example, the forwarding operation can be an operation where the user drags the edited content (or a file containing the edited content, such as a document, etc.) displayed by the first application to the application interface of the second application.

[0279] 2) The second application can obtain the edited content indicated by the user operation in response to the received user operation.

[0280] For example, the user can copy the edited content (or a file containing the edited content, such as a document, etc.) from the first application and paste the edited content into the second application so that the second application obtains the edited content. In this example, the user operation received by the second application can be the operation where the user pastes the original content into the second application.

[0281] In some embodiments of the present application, after receiving the edited content, the second application can display the edited content. After converting the edited content into multitasking prompt content, the second application can display the multitasking prompt content. Optionally, when the second application displays the edited content, it can, in response to the received conversion operation, generate multitasking prompt content based on the edited content. Among them, the conversion operation is used to indicate converting the original content into multitasking prompt content. Exemplarily, the conversion operation can be an operation where the user inputs a conversion instruction, or can be an operation where the user acts on a conversion control displayed by the second application. Among them, the conversion control can be used to trigger converting the edited content into multitasking prompt content.

[0282] Exemplarily, taking the second application as a smart assistant application and the original content as the text content in the original document in the above embodiments, the first application can be a document application. As Figure 15 shown, the document application can display the edited content, and the smart assistant application can display an application interface for interacting with the user. After the document application displays the edited content, it can, in response to the operation where the user drags the document to which the edited content belongs to the application interface displayed by the smart assistant application, send the document to the smart assistant application. The smart assistant application can display the document in the user interface. And it can, in response to the user's conversion operation (for example, the operation where the user inputs an instruction of "optimize the document according to the annotations" in the application interface of the smart assistant application), convert the edited content in the document into multitasking prompt content and feedback the multitasking prompt content (for example Figure 15 the document optimization draft shown).

[0283] Method 3: The electronic device executes the above-mentioned step S501 through the first application, executes the above-mentioned step S502 through the second application, and executes the above-mentioned step S503 through the second application or the third application.

[0284] Among them, the first application can generate or obtain the original content and can also display the original content. The second application can be a system application or a third-party application. The third application can be a system application or a third-party application.

[0285] In a possible method, the first application can generate or obtain the original content and display the original content. After the first application displays the original content, it can forward the original content to the second application in response to the received forwarding operation. Among them, the specific forwarding method can refer to the method in which the first application forwards the edited content to the second application in the above-mentioned method 2, which will not be elaborated here.

[0286] In another possible method, the second application can obtain the original content indicated by the user operation in response to the received user operation. Among them, the specific obtaining method can refer to the method in which the second application obtains the edited content according to the user operation in the above-mentioned method 2, which will not be elaborated here.

[0287] After the second application obtains the original content based on any of the above possible methods, it can display the original content. And it can generate the edited content and display the edited content in response to the operation of the user editing the original content.

[0288] In a possible method, after the second application displays the edited content, it can generate multitasking prompt content according to the edited content and display the multitasking prompt content in response to the received conversion operation. Among them, the specific implementation manner of this method can refer to the above-mentioned method 1, which will not be elaborated here.

[0289] In another possible method, after the second application displays the edited content, the third application can convert the edited content into multitasking prompt content in response to the received conversion operation after obtaining the edited content. Among them, the specific implementation manner of this method can refer to the above-mentioned method 2, which will not be elaborated here.

[0290] Exemplarily, taking the second application as a note application and the original content as the content in the original document described in the above embodiment as an example, the first application can be a document application. After the note application obtains and displays the original content, the user can edit the original content displayed by the note application in a handwriting manner, and then obtain Figure 16 the edited content shown in

[0291] It should be noted that the applications mentioned in the above methods can also be replaced by service processes or functional units, etc.

[0292] Based on the above method, the electronic device supports content editing and multitasking through various interaction methods. In practical applications, an appropriate method can be selected according to the specific scenario, so it has high flexibility and practicability.

[0293] Next, the training method of the AI model described in the embodiments of the present application will be described.

[0294] A possible training method of the AI model provided in the embodiments of the present application may include:

[0295] Step 1: The electronic device obtains a training data set.

[0296] In this step, the training data set obtained by the electronic device may include single-task prompt content and corresponding single-task processing results, and / or multi-task prompt content and corresponding multi-task processing results. Among them, the formats and specific contents of data such as single-task prompt content, single-task processing results, multi-task prompt content, and corresponding multi-task processing results can refer to the descriptions in the foregoing embodiments and will not be elaborated here.

[0297] After obtaining the single-task prompt content, the electronic device can input the single-task prompt content into the deep learning model for processing to obtain the output result corresponding to the single-task prompt content. After obtaining the multi-task prompt content, the electronic device can split the multi-task prompt content into multiple single-task prompt contents and input each single-task prompt content into the deep learning model for processing to obtain the output result corresponding to each single-task prompt content. For each single-task prompt content, the electronic device can use the output result corresponding to the single-task prompt content as the single-task processing result corresponding to the single-task prompt content, or can, in response to the received operation, use the modified output result as the single-task processing result corresponding to the single-task prompt content. Wherein, the operation is used to indicate modifying the output result.

[0298] For the multi-task prompt content, after the electronic device determines the single-task processing results corresponding to the multiple single-task prompt contents in the multi-task prompt content, it can summarize and integrate the task processing results corresponding to the multiple single-task prompt contents to obtain the multi-task processing result corresponding to the multi-task prompt content, and further obtain the multi-task prompt content and the corresponding multi-task processing result used as the training data set.

[0299] Optionally, the deep learning model can be an AI model with a set network structure or a set type. For example, the deep learning model can be an existing generative pre-trained transformer (GPT) model, large language model, etc.

[0300] Step 2: The electronic device trains the deep learning model using the training dataset to obtain the trained AI model.

[0301] In a possible solution, the electronic device may input the multi-task prompt content in the training dataset into a set model, and use the set model to generate corresponding processing results. And based on the processing results of the set model and the single-task processing results and / or multi-task processing results in the training data, perform reinforcement learning from human feedback (RLFH) on the set model, thereby optimizing the model parameters and processing effects of the set model, so that the set model learns the ability to process tasks in combination with the context and process multiple tasks in a single input sequence. Optionally, the set model may be an AI model.

[0302] The model trained based on the above method can be used as a model capable of recognizing and processing multi-task prompt content.

[0303] Of course, the electronic device may also use other training methods to train the set model, thereby obtaining an AI model capable of processing tasks in combination with the context and recognizing and processing multi-task prompt content. In the embodiments of the present application, no specific limitation is imposed on the specific model training method.

[0304] It should be noted that the implementation processes provided in the above embodiments are only examples of the applicable method processes of the embodiments of the present application. The execution order of each step can be adjusted accordingly according to actual needs, and other steps can also be added or some steps can be reduced.

[0305] Based on the above embodiments and the same technical concept, the embodiments of the present application also provide a content processing method, as Figure 17 shown in, this method may include:

[0306] S1701: The electronic device obtains first content, first location information, and first task information; wherein, the first location information is used to indicate the first location in the first content, and the first task information is used to indicate performing a first processing task on the first location.

[0307] Optionally, the electronic device may be the Figure 4a electronic device shown in, and the steps in this method may be executed by the content processing system deployed in the Figure 4a electronic device shown in. The specific execution method may refer to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0308] The following details the first content, the first location information, and the first task information.

[0309] In the first possible solution, the first content may include a first sub-content, the first position may be a position before or after the first sub-content, the first reference content may include the first sub-content, and the first reference content may be used as the context content of the first position.

[0310] Exemplarily, the first content may be the original content described in the foregoing embodiments, the first position may be any position in the original content described in the foregoing embodiments, the first sub-content may be the content before or after this position, and the first task information may be the task indication information corresponding to this position (or the position information of this position).

[0311] Example 8. In one example, the first content may be Figure 10 the text content shown in Figure 10 or Figure 11 "You have to be as tough as building the Great Wall!" shown in the (a) schematic diagram of Figure 11 The first position may be the position before the first sub-content, that is, the position corresponding to the editing opinion "Briefly introduce the project" shown in the (a) schematic diagram of

[0312] Optionally, if the first position is the position before the first sub-content, then in the content input to the first model, the first position information and the first task information are before the first sub-content; if the first position is the position after the first sub-content, then in the content input to the first model, the first position information and the first task information are after the first sub-content.

[0313] Exemplarily, the content input to the first model may be the multi-task prompt content described in the foregoing embodiments.

[0314] Example 9. In one example, based on the above Example 8, the content input to the first model may be Figure 11 the multi-task prompt text shown in the (b) schematic diagram of

[0315] In the second possible solution, on the basis of the first possible solution described above, the first content may further include a second sub-content, the first position may be the position between the first sub-content and the second sub-content, and the first reference content may further include the second sub-content. Optionally, in the content input to the first model, the first position information and the first task information are between the first sub-content and the second sub-content.

[0316] Example 10. In one example, the first content may be Figure 10 The text content shown in , the first sub - content can be Figure 10 or Figure 11 The content shown in the (a) schematic diagram in , "Hello everyone! I'm here today to tell you about the project we're about to launch! As we all know, this project is a huge undertaking,". The second sub - content can be Figure 10 or Figure 11 The content shown in the (a) schematic diagram in , "It has to be as tough as building the Great Wall!". The first position can be the position between the first sub - content and the second sub - content, that is, Figure 11 The position corresponding to the editorial comment "Briefly introduce the project" shown in the (a) schematic diagram in . Correspondingly, the content input to the first model can be Figure 11 The multi - task prompt text shown in the (b) schematic diagram in . Among them, the first position information includes two [ / Prmpt<0>], and the first task information is <Briefly introduce the project>. In this multi - task prompt text, the first position information and the first task information are located between the first sub - content and the second sub - content.

[0317] In the third possible scenario, the first content can include a third sub - content, and the third sub - content is located at the first position. Or, based on the above - mentioned first possible scenario or the above - mentioned second possible scenario, the first content can also include a third sub - content, and the third sub - content is located at the first position (that is, the first position is the position where the third sub - content is located). In this scenario, the second content specifically includes the content obtained by the first model performing the first processing task on the third sub - content according to the first reference content. In this scenario, when there is substantial content (i.e., the content is not empty) at the first position in the first content, the content obtained by the first model performing the first processing task on the first position according to the first reference content, that is, the content obtained by the first model performing the first processing task on the substantial content at the first position according to the first reference content.

[0318] Example 11. In an example, the first content can be Figure 10 The text content shown in , and the third sub - content can be Figure 10 The content shown in , "Just like the lyrics of the song 'We Are Different': 'We are different, our dreams are different'". The position where this third sub - content is located is the first position. The first sub - content can be Figure 10 Some or all of the text content before this third sub - content shown in , and the second sub - content can be Figure 10 Some or all of the text content after this third sub - content shown in .

[0319] In some embodiments of the present application, the first position information may include start position information and end position information; wherein, the start position information is used to indicate the starting point of the first position, and the end position information is used to indicate the ending point of the first position; the third sub - content is located between the starting point and the ending point in the first content.

[0320] For example, in the above Example 9 or Example 10, the start position information in the first position information may be Figure 11 the first one shown in the schematic diagram (b) in [ / Prmpt<0>](i.e., before [ / Prmpt<0>]<Briefly introduce the project>), and the end position information in the first position information may be Figure 11 the second one shown in the schematic diagram (b) in [ / Prmpt<0>](i.e., after [ / Prmpt<0>]<Briefly introduce the project>).

[0321] Optionally, in the content input to the first model, the third sub - content may be between the start position information and the end position information; and / or, in the content input to the first model, the first task information may be before the start position information, or between the start position information and the third sub - content, or between the third sub - content and the end position information, or after the end position information.

[0322] For example, in the above Example 11, the content input to the first model is Figure 11 the multi - task prompt text shown in the schematic diagram (b) in [ / Prmpt<0>]. In this multi - task prompt text, the third sub - content (i.e., just like the lyrics of the song "We Are Different": "We are different, our dreams are different") is located between the start position information (e.g., [ / Prmpt<4>) and the end position information (e.g., [ / Prmpt<4>]). The first task information (e.g., <Play a different inspiring song>]) is between the start position information and the third sub - content.

[0323] In some embodiments of the present application, the first content may include at least one of the following: text, image, audio, video, code.

[0324] Among them, in a possible case, when the first content is the first image, the first position information is the semantic mask image corresponding to the content at the first position in the first image. For example, the first content (the first image) may be Figure 12 the original image shown in [ / Prmpt<0>], the first position may be Figure 12 the position where the background is located in [ / Prmpt<0>], the first position information may be Figure 12 the mask region 1 in the semantic mask image shown in [ / Prmpt<0>], and the first task information may be Figure 12 the editing opinion 1 shown in [ / Prmpt<0>]. Again, for example, the first content (the first image) may be Figure 12 In the original image shown in, the first position can be Figure 12 the position where the person in the original image shown in is located, and the first position information can be Figure 12 the mask region 2 in the semantic mask image shown in, and the first task information can be Figure 12 the edit opinion 2 shown in.

[0325] In the embodiments of the present application, the semantic mask image refers to a specific form of image output generated by a semantic segmentation model for an original image, where each pixel in the semantic mask image is labeled as one of the predefined categories. The semantic segmentation model can generate a multi-label feature map (i.e., the semantic mask image) based on the original image, where the value of each pixel in the feature map represents the label of a different object category. Based on the semantic segmentation model, objects can be accurately classified and precise boundaries can be drawn at a fine pixel level.

[0326] The method for obtaining the first content, the first position information, and the first task information will be described in detail below.

[0327] In one possible solution, the electronic device can be a server, that is, the content display method provided in the embodiments of the present application can be applied to the server. Optionally, the server can be a cloud server. In this scenario, the electronic device can receive the first content, the first position information, and the first task information from the terminal device, and thus obtain the first content, the first position information, and the first task information. Among them, the method for the terminal device to obtain the first content, the first position information, and the first task information can refer to the obtaining method when the electronic device is a terminal device in the foregoing or the following text, which will not be elaborated here for the time being.

[0328] In another possible solution, the electronic device can be a terminal device, that is, the content display method provided in the embodiments of the present application can be applied to the terminal device. In this scenario, the electronic device can first obtain the first content and the first indication information, and then determine the first position information and the first task information according to the first content and the first indication information, and thus obtain the first content, the first position information, and the first task information. Among them, the first indication information is used to indicate to perform the first processing task on the first position in the first content.

[0329] Exemplarily, the first content can be the original content described in the foregoing embodiments, the first indication information can be the task indication information in the edited content described in the foregoing embodiments, and the first content and the first indication information can form the edited content described in the foregoing embodiments.

[0330] As an alternative implementation, after the electronic device obtains the first content and the first indication information, before determining the first location information and the first task information according to the first content and the first indication information, the electronic device may display the first content and the first indication information and may receive a first operation. The first operation is used to indicate: according to the first content, convert the first indication information into the first location information and the first task information. The electronic device may, in response to the received first operation, determine the first location information and the first task information according to the first content and the first indication information. Optionally, after the electronic device determines the first location information and the first task information according to the first content and the first indication information, the electronic device may display the first content, the first location information, and the first task information so that the user can view the content to be input into the first model after conversion.

[0331] In some embodiments of the present application, the electronic device may obtain the first content and the first indication information in any of the following ways:

[0332] 1) Obtain and display the first content through a first application, and determine the first indication information according to a second operation acting on the first content received.

[0333] The first application is an application installed in the electronic device. The first application may be used to generate the first content and may determine the first indication information according to an editing operation (i.e., the second operation) of the user on the first content.

[0334] Exemplarily, the first application may be the first application described in Method 1 in the content conversion control method in the foregoing embodiments, and the second operation may be the operation of editing the original content described in Method 1 in the content conversion control method in the foregoing embodiments.

[0335] For the specific implementation process of this method, reference may be made to Method 1 in the content conversion control method described in the foregoing embodiments, and details are not described herein again.

[0336] 2) Obtain the first content and the first indication information indicated by a third operation received through the first application.

[0337] Exemplarily, the first application may be the second application described in Method 2 in the content conversion control method in the foregoing embodiments, and the second operation may be the forwarding operation or the user operation described in Method 2 in the content conversion control method in the foregoing embodiments.

[0338] For the specific implementation process of this method, reference may be made to Method 2 in the content conversion control method described in the foregoing embodiments, and details are not described herein again.

[0339] 3) Receive, by a first application, first content and first indication information sent by a second application; wherein, the second application is configured to generate the first content and the first indication information.

[0340] In one example, the first application may be the second application described in Method 2 of the content conversion control method in the foregoing embodiments, the second application may be the first application described in Method 2 of the content conversion control method in the foregoing embodiments, and the second operation may be the forwarding operation described in Method 2 of the content conversion control method in the foregoing embodiments. The specific implementation process of this method may be implemented with reference to Method 2 of the content conversion control method described in the foregoing embodiments, and will not be elaborated herein.

[0341] In another example, the first application may be the third application described in Method 3 of the content conversion control method in the foregoing embodiments, and the second application may be the second application described in Method 3 of the content conversion control method in the foregoing embodiments. The specific implementation process of this method may be implemented with reference to Method 3 of the content conversion control method described in the foregoing embodiments, and will not be elaborated herein.

[0342] S1702: The electronic device inputs the first content, the first location information, and the first task information into a first model to determine second content; wherein, the second content includes: content obtained by the first model performing a first processing task on a first location according to first reference content; the first reference content includes part or all of the first content except the first location.

[0343] In some embodiments of the present application, the first model belongs to a generative artificial intelligence AI model.

[0344] Exemplarily, the first model may be the AI model described in the foregoing embodiments.

[0345] The above method provides a content processing method in a single-task scenario (i.e., there is a set of location information and task information for the first content (the set of location information is the above-mentioned first location information and first task information)). In a multi-task scenario, there may be multiple sets of location information and task information for the first content. Wherein, each set of location information and task information includes location information for indicating a location in the first content and task information for indicating a processing task to be performed on that location. In a multi-task scenario, the electronic device may use the first model to uniformly process the first content and the corresponding multiple sets of location information and task information, and then obtain the task processing results corresponding to each set of location information and task information.

[0346] Based on this, in a possible scenario, the electronic device may also refer to the above method for obtaining the first position information and the second position information to obtain the second position information and the second task information; wherein, the second position information is used to indicate the second position in the first content, and the second task information is used to indicate the execution of the second processing task for the second position. When the electronic device inputs the first content, the first position information, and the first task information into the first model, it may simultaneously input the second position information and the second task information into the first model. That is to say, the electronic device may simultaneously input the first content, the first position information, the first task information, the second position information, and the second task information into the first model for processing. Based on this, the second content obtained by the first model processing includes both the content obtained by the first model according to the first reference content and performing the first processing task for the first position as described above, and the content obtained by the first model according to the second reference content and performing the second processing task for the second position; wherein, the second reference content includes part or all of the content in the first content except the second position. Of course, the content input into the first model may include more groups of position information and task information, which may be specifically implemented with reference to the above method and will not be enumerated one by one in the embodiments of the present application.

[0347] Among them, regarding the second position information, the second task information, and the second reference information, etc., reference may be respectively made to the descriptions of the first position information, the first task information, and the second reference information, etc. in the above text, and details will not be elaborated here one by one.

[0348] The training process of the first model will be described in detail below.

[0349] The training process of the first model may be executed before step S1702. In some embodiments of the present application, the training process of the first model may include the following steps 1 to 2.

[0350] Step 1: The electronic device obtains training data; wherein, the training data includes the third content, the third position information, the third task information, and the fourth content; wherein, the third position information is used to indicate the third position in the third content, and the third task information is used to indicate the execution of the third processing task for the third position; the fourth content includes: the content obtained by performing the third processing task for the third position according to part or all of the content in the third content except the third position.

[0351] In some embodiments of the present application, when the first model is used in a single-task scenario, the training data may include the original content (i.e., the above-mentioned third content), a set of position information and task information (i.e., the above-mentioned third position information and third task information), and the corresponding task processing result (i.e., the above-mentioned fourth content).

[0352] In some embodiments of the present application, when the first model is used in a multi-task scenario, the training data may include the original content, multiple sets of location information and task information, and the corresponding task processing results. Based on this, in a possible scenario, in addition to the above-mentioned third content (as the original content), the third location information, and the third task information (as a set of location information and task information), the training data may further include the fourth location information and the fourth task information (for use as another set of location information and task information), and the fourth content further includes: the content obtained by performing a fourth processing task on the fourth location according to some or all of the content in the third content other than the fourth location; wherein, the fourth location information is used to indicate the fourth location in the third content, and the fourth task information is used to indicate performing the fourth processing task on the fourth location. Of course, the training data may include more sets of location information and task information and the corresponding task processing results, and the specific implementation may refer to the above method, which will not be listed one by one in the embodiments of the present application.

[0353] In some embodiments of the present application, in the above-mentioned single-task scenario, the electronic device may obtain the training data according to the method described in the following steps A1 to A4.

[0354] A1: The electronic device obtains the third content and the second indication information; wherein, the second indication information is used to indicate performing a third processing task on the third location in the third content.

[0355] A2: The electronic device determines the third location information and the third task information according to the second indication information.

[0356] A3: The electronic device inputs the third content and the second indication information into the second model to determine the fifth content; or inputs the content at the third location in the third content and the second indication information into the second model to determine the fifth content; or when there is no substantial content at the third location, inputs the second indication information into the second model to determine the fifth content; wherein, the fifth content includes the content obtained by the second model performing the third processing task according to the input content.

[0357] Among them, the second model may be a generative AI model used in the traditional solution, such as an existing generative AI model, a large language model, etc.

[0358] Regarding the above-mentioned third content, second indication information, third location information, and third task information, reference may be made to the first content, first indication information, first location information, and first task information described in the foregoing embodiments respectively, and details will not be repeated here.

[0359] A4: The electronic device determines the fourth content according to the fifth content.

[0360] Among them, as an optional implementation, the electronic device may use the fifth content as the fourth content.

[0361] As another optional implementation, the electronic device may modify the fifth content in response to a received modification operation to obtain the fourth content; wherein, the modification operation is used to indicate modifying the fifth content to the fourth content. Based on this method, the fourth content obtained by processing with the second model can be corrected to obtain the fifth content by manual modification, which can ensure the accuracy of the task processing result. When using the fifth content as training data, the accuracy of the training data can be ensured, thereby improving the model training effect.

[0362] In some embodiments of the present application, in the above multi-task scenario, the electronic device may obtain training data according to the method described in the following steps B1 to B4.

[0363] B1: The electronic device obtains the third content, the second indication information, and the third indication information; wherein, the second indication information is used to indicate performing a third processing task on a third position in the third content; the third indication information is used to indicate performing a fourth processing task on a fourth position in the third content.

[0364] B2: The electronic device inputs the third content and the second indication information into the second model to determine the fifth content; or inputs the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or inputs the second indication information into the second model to determine the fifth content; wherein, the fifth content includes the content obtained by the second model performing the third processing task according to the input content.

[0365] B3: The electronic device inputs the third content and the third indication information into the second model to determine the sixth content; or inputs the content at the fourth position in the third content and the third indication information into the second model to determine the sixth content; or inputs the third indication information into the second model to determine the sixth content; wherein, the sixth content includes the content obtained by the second model performing the fourth processing task according to the input content.

[0366] The above steps B1 to B3 may be implemented by referring to the above steps A1 or A2 respectively, and will not be elaborated here one by one.

[0367] B4: The electronic device determines the fourth content according to the fifth content and the sixth content.

[0368] Specifically, the electronic device may use the content obtained by combining the seventh content determined according to the fifth content and the eighth content determined according to the sixth content as the fourth content. The seventh content is the fifth content or the content obtained by modifying the fifth content in response to the received operation; the eighth content is the sixth content or the content obtained by modifying the sixth content in response to the received operation. Specifically, it can be implemented with reference to the method of determining the fourth content according to the fifth content above, which will not be elaborated here.

[0369] Step 2: The electronic device trains the set model according to the training data to obtain the first model.

[0370] Specifically, the electronic device may input the training data obtained based on the above method into the set model to determine the ninth content. In the above single-task scenario, the ninth content includes the content obtained by the set model performing the third processing task on the third position in the third content. In the above multi-task scenario, the ninth content further includes the content obtained by the set model performing the fourth processing task on the fourth position in the third content. After determining the ninth content, the electronic device may adjust the set model according to the third content and the ninth content to obtain the first model.

[0371] Exemplarily, the third content may be the single-task processing result / multi-task processing result in the training data described in the training method of the AI model in the foregoing embodiments, and the ninth content may be the processing result of the set model described in the training method of the AI model in the foregoing embodiments.

[0372] In the above method, the specific steps executed by the electronic device may also refer to the relevant introduction in the foregoing embodiments, and will not be elaborated here.

[0373] Based on the above embodiments and the same technical concept, an embodiment of the present application further provides a content processing method, which can be applied to a content processing system including a server and a terminal device, as Figure 18 shown in, this method may include:

[0374] S1801: The terminal device sends the first content, the first position information, and the first task information to the server; wherein, the first position information is used to indicate the first position in the first content, and the first task information is used to indicate performing the first processing task on the first position.

[0375] S1802: The server receives the first content, the first position information, and the first task information from the terminal device.

[0376] S1803: The server inputs the first content, the first location information, and the first task information into the first model to determine the second content, where the second content includes the content obtained by the first model performing the first processing task for the first location according to the first reference content, and the first reference content includes some or all of the content in the first content except the first location.

[0377] S1804: The server sends the second content to the terminal device.

[0378] S1805: The terminal device receives the second content from the server.

[0379] Optionally, the terminal device may be Figure 4b the terminal device shown in Figure 4b and the server may be Figure 4b the server shown in. The steps in this method may be executed by

[0380] the content processing system composed of the terminal device and the server shown in Figure 17 and the specific execution method may refer to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0381] In the above method, for the specific steps executed by the terminal device or the server, reference may be made to Figure 17 the execution method of the corresponding steps in the method shown, and no more details will be given here.

[0381] Based on the above embodiments and the same technical concept, an embodiment of the present application further provides a content processing system, which may include a server and a terminal device. The terminal device is configured to: send the first content, the first location information, and the first task information to the server; where the first location information is used to indicate the first location in the first content, and the first task information is used to indicate performing the first processing task for the first location; the server is configured to: receive the first content, the first location information, and the first task information from the terminal device; input the first content, the first location information, and the first task information into the first model to determine the second content, where the second content includes the content obtained by the first model performing the first processing task for the first location according to the first reference content, and the first reference content includes some or all of the content in the first content except the first location; send the second content to the terminal device; the terminal device is further configured to: receive the second content from the server.

[0382] Among them, for the specific functions or execution methods of the terminal device and the server, reference may be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0383] Based on the above embodiments and the same technical concept, an embodiment of the present application further provides an electronic device, which is used to implement the content processing method applied to the electronic device, the server, or the terminal device provided in the embodiments of the present application. As Figure 19 As shown in the figure, the electronic device 1900 may include: a memory 1901, one or more processors 1902, and one or more computer programs (not shown in the figure). Each of the above components may be coupled through one or more communication buses 1903. Optionally, the electronic device 1900 may further include a display screen 1904.

[0384] Among them, one or more computer programs (codes) are stored in the memory 1901, and the one or more computer programs include computer instructions; the one or more processors 1902 call the computer instructions stored in the memory 1901, so that the electronic device 1900 executes the content processing method provided in the above embodiments of the present application applied to the electronic device or the server or the terminal device.

[0385] In a specific implementation, the memory 1901 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 1901 may store an operating system (hereinafter referred to as the system), such as an embedded operating system like ANDROID, IOS, WINDOWS, or LINUX. The memory 1901 may be used to store the implementation program of the embodiments of the present application. The memory 1901 may also store a network communication program, which may be used to communicate with one or more additional devices, one or more user devices, and one or more network devices.

[0386] One or more processors 1902 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application solution.

[0387] The display screen 1904 is used to display relevant user interfaces such as application interfaces.

[0388] It should be noted that Figure 19 This is only one implementation manner of the electronic device 1900 provided in the embodiments of the present application. In actual applications, the electronic device 1900 may further include more or fewer components. Specifically, reference may be made to Figure 3 the specific structure and description shown, which are not limited here.

[0389] Based on the above embodiments and the same technical concept, the embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a computer, the computer executes the method provided in the above embodiments applied to the electronic device or the server or the terminal device.

[0390] Based on the above embodiments and the same inventive concept, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions run on a computer, the computer is caused to execute the method applied to an electronic device, a server, or a terminal device provided in the above embodiments.

[0391] In the method provided by the embodiment of the present application, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as an SSD), etc.

[0392] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.< / id> < / id> < / id> < / id>

Claims

1. A content processing method, characterized in that, it includes: Obtain a first content, first position information, and first task information; wherein, the first position information is used to indicate a first position in the first content, and the first task information is used to indicate to perform a first processing task on the first position; Input the first content, the first position information, and the first task information into a first model to determine a second content; wherein, the second content includes: the content obtained by the first model performing the first processing task on the first position according to a first reference content; the first reference content includes part or all of the content in the first content except the first position.

2. The method according to claim 1, characterized in that, the first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content.

3. The method according to claim 2, characterized in that, the first content further includes a second sub-content, the first position is a position between the first sub-content and the second sub-content, and the first reference content further includes the second sub-content.

4. The method according to any one of claims 1 to 3, characterized in that, the first content includes a third sub-content, and the third sub-content is located at the first position; the second content specifically includes the content obtained by the first model performing the first processing task on the third sub-content according to the first reference content.

5. The method according to claim 4, characterized in that, the first position information includes start position information and end position information; wherein, the start position information is used to indicate the start point of the first position, and the end position information is used to indicate the end point of the first position; the third sub-content is located between the start point and the end point in the first content.

6. The method according to claim 5, characterized in that, in the content input to the first model, the third sub-content is between the start position information and the end position information; and / or in the content input to the first model, the first task information is before the start position information, or between the start position information and the third sub-content, or between the third sub-content and the end position information, or after the end position information.

7. The method according to any one of claims 1 to 6, characterized in that, the method is applied to a server; the obtaining of the first content, the first position information, and the first task information includes: Receiving the first content, the first position information, and the first task information from a terminal device.

8. The method according to any one of claims 1 to 6, characterized in that, the obtaining of the first content, the first position information, and the first task information includes: Obtaining the first content and first indication information; wherein, the first indication information is used to indicate to perform the first processing task on the first position in the first content. Determine the first location information and the first task information according to the first content and the first indication information.

9. The method according to claim 8, wherein, after obtaining the first content and the first indication information and before determining the first location information and the first task information according to the first content and the first indication information, the method further includes: displaying the first content and the first indication information; receiving a first operation; wherein the first operation is used to indicate: converting the first indication information into the first location information and the first task information according to the first content.

10. The method according to claim 8 or 9, wherein, the obtaining of the first content and the first indication information includes: obtaining and displaying the first content through a first application, and determining the first indication information according to a second operation acting on the first content; or obtaining the first content and the first indication information indicated by a third operation received through a first application; or receiving the first content and the first indication information sent by a second application through a first application; wherein the second application is used to generate the first content and the first indication information.

11. The method according to any one of claims 8 to 10, wherein, the method is applied to a terminal device.

12. The method according to any one of claims 1 to 11, wherein, before inputting the first content, the first location information and the first task information into a first model to determine a second content, the method further includes: obtaining second location information and second task information; wherein the second location information is used to indicate a second location in the first content, and the second task information is used to indicate performing a second processing task on the second location; when inputting the first content, the first location information and the first task information into the first model, the method further includes: inputting the second location information and the second task information into the first model; the second content further includes: content obtained by the first model performing the second processing task on the second location according to second reference content; wherein the second reference content includes part or all of the first content except the second location.

13. The method according to any one of claims 1 to 12, wherein, before inputting the first content, the first location information and the first task information into a first model to determine a second content, the method further includes: obtaining training data; wherein the training data includes third content, third location information, third task information and fourth content; wherein the third location information is used to indicate a third location in the third content, and the third task information is used to indicate performing a third processing task on the third location; the fourth content includes: content obtained by performing the third processing task on the third location according to part or all of the third content except the third location. Training the set model according to the training data to obtain the first model.

14. The method according to claim 13, wherein, the training data further includes fourth position information and fourth task information; wherein, the fourth position information is used to indicate a fourth position in the third content, and the fourth task information is used to indicate a fourth processing task to be performed on the fourth position; the fourth content further includes: content obtained by performing the fourth processing task on the fourth position according to part or all of the content in the third content other than the fourth position.

15. The method according to any one of claims 1 to 14, wherein, the first content includes at least one of the following: text, image, audio, video, code.

16. The method according to any one of claims 1 to 15, wherein, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.

17. A content processing system, wherein, comprising: a terminal device and a server; the terminal device is configured to: send the first content, the first position information, and the first task information to the server; wherein, the first position information is used to indicate a first position in the first content, and the first task information is used to indicate a first processing task to be performed on the first position; the server is configured to: receive the first content, the first position information, and the first task information from the terminal device; input the first content, the first position information, and the first task information into a first model to determine a second content, wherein the second content includes content obtained by the first model performing the first processing task on the first position according to a first reference content, and the first reference content includes part or all of the content in the first content other than the first position; send the second content to the terminal device; the terminal device is further configured to: receive the second content from the server.

18. The system according to claim 17, wherein, the first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content.

19. The system according to claim 18, wherein, the first content further includes a second sub-content, the first position is a position between the first sub-content and the second sub-content, and the first reference content further includes the second sub-content.

20. The system according to any one of claims 17 to 19, wherein, the first content includes a third sub-content, and the third sub-content is located at the first position; the second content specifically includes content obtained by the first model performing the first processing task on the third sub-content according to the first reference content.

21. The system according to any one of claims 17 to 20, wherein, the terminal device is further configured to: Before sending the first content, the first location information, and the first task information to the server, obtain the first content and the first indication information; wherein, the first indication information is used to indicate to perform the first processing task on the first location in the first content. Determine the first location information and the first task information according to the first content and the first indication information.

22. The system according to claim 21, wherein, the terminal device is further configured to: after obtaining the first content and the first indication information, and before determining the first location information and the first task information according to the first content and the first indication information, display the first content and the first indication information, and receive a first operation; wherein, the first operation is used to indicate: according to the first content, convert the first indication information into the first location information and the first task information.

23. The system according to claim 21 or 22, wherein, when the terminal device obtains the first content and the first indication information, it is specifically configured to: obtain and display the first content through a first application, and determine the first indication information according to a second operation acting on the first content received; or obtain the first content and the first indication information indicated by a third operation received through a first application; or receive the first content and the first indication information sent by a second application through a first application; wherein, the second application is used to generate the first content and the first indication information.

24. The system according to any one of claims 17 to 23, wherein, the terminal device is further configured to: receive a third operation before sending the first content, the first location information, and the first task information to the server; wherein, the third operation is used to indicate to perform a processing task according to the first content, the first location information, and the first task information.

25. The system according to any one of claims 17 to 24, wherein, the terminal device is further configured to: send second location information and second task information to the server; wherein, the second location information is used to indicate a second location in the first content, and the second task information is used to indicate to perform a second processing task on the second location; the server is further configured to: when inputting the first content, the first location information, and the first task information into a first model, input the second location information and the second task information into the first model; the second content further includes: the content obtained by the first model performing the second processing task on the second location according to second reference content; wherein, the second reference content includes part or all of the first content except the second location.

26. The system according to any one of claims 17 to 25, wherein, the server is further configured to: Before inputting the first content, the first location information, and the first task information into a first model to determine the second content, training data is obtained, and a set model is trained according to the training data to obtain the first model; wherein the training data includes third content, third location information, third task information, and fourth content; wherein the third location information is used to indicate a third location in the third content, and the third task information is used to indicate performing a third processing task on the third location; the fourth content includes: content obtained by performing the third processing task on the third location according to some or all of the content in the third content other than the third location.

27. The system according to claim 26, wherein, the training data further includes fourth location information and fourth task information; wherein the fourth location information is used to indicate a fourth location in the third content, and the fourth task information is used to indicate performing a fourth processing task on the fourth location; the fourth content further includes: content obtained by performing the fourth processing task on the fourth location according to some or all of the content in the third content other than the fourth location.

28. An electronic device, wherein, the electronic device includes a memory and one or more processors; wherein the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the one or more processors, the electronic device is caused to execute the method according to any one of claims 1 to 16.

29. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program runs on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 16.

Citation Information

Cited By

  • Content processing method, system and electronic device

    EP4749511A1