Content processing method, system and electronic device
By using the context information of the content in the AI model for task processing, the instability problem of the AI model in complex task scenarios is solved, and the accuracy and practicality of task processing are improved.
Patent Information
- Application Number
- PCT/CN2024/136112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-19
AI Technical Summary
The existing AI models have unstable task execution effects, especially in complex task scenarios, which are difficult to meet task processing requirements, resulting in low practicality.
By obtaining content, location information and task information, input it into the AI model to determine more accurate processing results, and using the context information in the content for task processing.
The accuracy and practicality of the AI model in task processing is improved, making the output results easier to meet task processing requirements.
Smart Images

Figure CN2024136112_19062025_PF_FP_ABST
Abstract
Description
Content processing method, system and electronic device
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 12, 2023, with application number 202311705791.3 and application name "A content processing method, system and electronic device", all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the technical field of electronic devices, and in particular to a content processing method, system and electronic device. Background Art
[0004] An artificial intelligence (AI) model (also known as an AI big model or an AI pre-trained big model) is a model that can adapt to a series of tasks after being trained with large-scale data. The AI model is a type of machine learning model. Currently, some AI models (such as generative AI models, large language models, etc.) can perform corresponding tasks based on the input task instruction information and output task processing results, or the AI model can perform corresponding tasks for the content based on the input content and the task instruction information corresponding to the content and output task processing results. However, in actual applications, the task execution effect of the AI model may be unstable. In some scenarios (such as complex task scenarios, etc.), the task processing results output by the AI model may be difficult to meet the task processing requirements, resulting in the low practicality of the AI model. Summary of the Invention
[0005] The present application provides a content processing method, system and electronic device to improve the effect of processing tasks through models, thereby improving the practicality of the models.
[0006] In a first aspect, an embodiment of the present application provides a content processing method, the method comprising: obtaining first content, first location information, and first task information; wherein the first location information is used to indicate a first location in the first content, and the first task information is used to indicate the execution of a first processing task for the first location; inputting the first content, the first location information, and the first task information into a first model to determine a second content; wherein the second content comprises: the content obtained by the first model executing the first processing task for the first location based on the first reference content; the first reference content comprises part or all of the content in the first content except the first location. Optionally, the first model is an AI model.
[0007] In this method, in the first content, the first reference content can be used as the context content of the first position. By inputting the first content, the first position information and the first task information into the first model, the electronic device can enable the first model to perform task processing on the first position with reference to the context content of the first position when executing the first processing task indicated by the first task information for the first position indicated by the first position information, thereby obtaining a more accurate processing result. Therefore, the above method helps to improve the accuracy of the results output when the electronic device uses the first model to perform a task, so that the output results can more easily meet the task processing requirements, thereby improving the practicality of the first model. When the first model is an AI model, the above method can improve the effect of processing tasks through the AI model, thereby improving the practicality of the AI model.
[0008] In one possible design, the first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content. In this method, the contextual content of the first position can be the content of positions on either side of the first position, which provides greater flexibility and facilitates controlling the data volume of the contextual content of the first position, thereby ensuring a certain level of processing efficiency.
[0009] Optionally, if the first position is before the first sub-content, then in the content input into the first model, the first position information and the first task information are located before the first sub-content; if the first position is after the first sub-content, then in the content input into the first model, the first position information and the first task information are located after the first sub-content. In this method, the positional relationship between the position information / task information and the sub-content in the content input into the first model can be associated with the positional relationship between the position corresponding to the position information and the sub-content in the first content, which can improve the consistency of the content conversion process and, if the content input into the first model is displayed, can improve the readability of the displayed content input into the first model.
[0010] In one possible design, the first content also includes a second sub-content, the first position is between the first and second sub-contents, and the first reference content also includes the second sub-content. In this method, the contextual content of the first position can be derived from the content of positions on both sides of the first position, which provides greater comprehensiveness and helps improve the accuracy of task processing.
[0011] Optionally, in the content input into the first model, the first location information and the first task information are located between the first sub-content and the second sub-content. In this method, the positional relationship between the location information / task information and the sub-content in the content input into the first model can be associated with the positional relationship between the position corresponding to the location information and the sub-content in the first content. This can improve consistency during the content conversion process and, if the content input into the first model is displayed, can improve the readability of the displayed content input into the first model.
[0012] In one possible design, the first content includes a third sub-content, and the third sub-content is located at the first position; the second content specifically includes the content obtained by the first model performing the first processing task on the third sub-content based on the first reference content.
[0013] In conjunction with the above method, the first position in the first content may contain a third sub-content, or it may not contain a sub-content, i.e., the content is empty. Regardless of whether a sub-content exists at the first position in the first content, the electronic device can, based on the above method and the context of the first position, perform relatively accurate task processing for the first position. Therefore, the above method is more effective in performing task processing using the first model.
[0014] In one possible design, the first position information includes starting position information and ending position information; wherein, the starting position information is used to indicate the starting point of the first position, and the ending position information is used to indicate the ending point of the first position; the third sub-content is located between the starting point and the ending point in the first content.
[0015] In this method, the starting position information and the ending position information can clearly and accurately indicate the position of the third sub-content in the first sub-content, which facilitates the first model to identify the third sub-content and perform subsequent processing, thereby improving processing efficiency.
[0016] In one possible design, in the content input into the first model, the third sub-content is between the starting position information and the ending position information; and / or, in the content input into the first model, the first task information is before the starting position information, or between the starting position information and the third sub-content, or between the third sub-content and the ending position information, or after the ending position information.
[0017] In this method, the positional relationship between the start position information and the end position information in the position information and the sub-content, task information, etc. in the content input to the first model can be flexibly set, which can improve the flexibility and practicality of the solution.
[0018] In one possible design, the method is applied to a server; and obtaining the first content, the first location information, and the first task information includes: receiving the first content, the first location information, and the first task information from a terminal device.
[0019] When the content processing method provided above is applied to a server, the server can obtain the content required to be input into the first model, namely, the first content, the first location information, and the first task information, from the terminal device, and process the content using the first model. Therefore, the method can be applied in scenarios where the server uses the first model to perform task processing on content provided by the terminal device, thereby reducing the task processing workload on the terminal device and improving processing efficiency on the terminal device side.
[0020] In one possible design, obtaining the first content, first location information and first task information includes: obtaining the first content and first indication information; wherein the first indication information is used to indicate the execution of the first processing task for the first location in the first content; and determining the first location information and the first task information based on the first content and the first indication information.
[0021] The first indication information can be used by the electronic device to identify or determine the first location and the first processing task, and the first location information and the first task information can be used by the first model to identify or determine the first location and the first processing task, respectively. By determining the first location information and the first task information based on the first content and the first indication information, the location information and task information recognizable by the electronic device can be converted into location information and task information recognizable by the first model, so that the first model can perform corresponding processing, thereby improving the processing efficiency of the first model and the accuracy of task execution.
[0022] In one possible design, after obtaining the first content and the first indication information, and before determining the first location information and the first task information based on the first content and the first indication information, the method further includes: displaying the first content and the first indication information; receiving a first operation; wherein the first operation is used to indicate: converting the first indication information into the first location information and the first task information based on the first content.
[0023] In this method, the electronic device can trigger the conversion of the execution instruction information based on the operation performed by the user, that is, the determination of the position information and the task information, so as to facilitate the user to control the content processing process.
[0024] In one possible design, obtaining the first content and the first indication information includes: obtaining and displaying the first content through a first application, and determining the first indication information based on a received second operation acting on the first content; wherein the second operation is used to indicate the first indication information; or, obtaining, through the first application, the first content and the first indication information indicated by a received third operation; or, receiving, through the first application, the first content and the first indication information sent by a second application; wherein the second application is used to generate the first content and the first indication information.
[0025] Optionally, before receiving the first content and the first indication information sent by the second application through the first application, the method also includes: acquiring and displaying the first content through the second application, and determining the first indication information based on the received fourth operation acting on the first content.
[0026] Among them, the first application and the second application are applications installed in the electronic device. The electronic device can generate and display the first content through the first application, and the user can edit the first content displayed by the first application by performing the second operation, and the first application of the electronic device can determine the first indication information based on the second operation. Alternatively, the user can directly indicate the first content and the second indication information to the first application of the electronic device by performing a third operation such as a paste operation. Alternatively, the electronic device can generate the first content through the second application and determine the first indication information based on the user's operation on the first content, and can send the first content and the first indication information to the first application through the second application. In the above method, the electronic device can use at least one application to obtain the first content and the first indication information in different ways, which is highly flexible and practical.
[0027] Optionally, inputting the first content, the first location information, and the first task information into the first model to determine the second content includes: inputting the first content, the first location information, and the first task information into the first model via the first application to determine the second content. Optionally, before inputting the first content, the first location information, and the first task information into the first model via the first application to determine the second content, the method further includes: displaying the first content, the first location information, and the first task information via the first application; and receiving a fifth operation via the first application; wherein the fifth operation is used to instruct the execution of a processing task based on the first content, the first location information, and the first task information.
[0028] By displaying the content input to the first model, namely the first content, the first location information, and the first task information, the user can conveniently view the content input to the first model and trigger the execution of the content processing task based on the corresponding operation performed by the user. This facilitates the user to control the content processing process, thereby improving the observability and controllability of the content processing process.
[0029] In one possible design, the method is applied to a terminal device.
[0030] In one possible design, before inputting the first content, the first location information and the first task information into the first model to determine the second content, the method further includes: obtaining second location information and second task information; wherein, the second location information is used to indicate the second location in the first content, and the second task information is used to indicate the execution of a second processing task for the second location; when inputting the first content, the first location information and the first task information into the first model, the method further includes: inputting the second location information and the second task information into the first model; the second content also includes: the content obtained by the first model executing the second processing task for the second location based on the second reference content; wherein, the second reference content includes part or all of the content in the first content except the second location.
[0031] Based on the above method, the electronic device can use the first model to simultaneously process multiple groups of tasks for the first content, such as the first processing task and the second processing task, thereby achieving the effect of simultaneously processing multiple tasks using the same model. Therefore, the above method can be applied to scenarios where multiple tasks are processed simultaneously, and can improve the efficiency of multitasking while ensuring the processing effect of each task in the multitasking.
[0032] In one possible design, before inputting the first content, the first location information and the first task information into the first model to determine the second content, the method further includes: obtaining training data; wherein the training data includes third content, third location information, third task information and fourth content; wherein the third location information is used to indicate the third location in the third content, and the third task information is used to indicate the execution of a third processing task for the third location; the fourth content includes: content obtained by executing the third processing task for the third location based on part or all of the content in the third content except the third location; and the set model is trained according to the training data to obtain the first model.
[0033] Based on the above method, the electronic device can implement training of a first model for processing a single task.
[0034] In one possible design, the acquiring of training data includes: acquiring the third content and second indication information; wherein the second indication information is used to indicate the execution of the third processing task for the third position in the third content; determining the third position information and the third task information based on the second indication information; inputting the third content and the second indication information into the second model to determine the fifth content; or, inputting the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, when there is no content at the third position, inputting the second indication information into the second model to determine the fifth content; wherein the fifth content includes the content obtained by the second model executing the third processing task based on the input content; and determining the fourth content based on the fifth content.
[0035] Optionally, determining the fourth content based on the fifth content includes: using the fifth content as the fourth content, or, in response to a received sixth operation, modifying the fifth content to obtain the fourth content; wherein the sixth operation is used to instruct to modify the fifth content to the fourth content. The second model can be a set generative AI model.
[0036] The above method for obtaining training data can be applied to the scenario of training a first model for processing a single task. Through the above method, the electronic device can use the set second model to collect partial training data used as the task processing result. Among them, by modifying the content obtained by processing the second model, that is, the fifth content, according to user operations, the accuracy of the task processing result data in the training data can be improved through manual correction. It can further improve the accuracy of model training using training data, thereby obtaining a first model with better processing effect.
[0037] In one possible design, the training data also includes fourth position information and fourth task information; wherein, the fourth position information is used to indicate the fourth position in the third content, and the fourth task information is used to indicate the execution of a fourth processing task for the fourth position; the fourth content also includes: content obtained by executing the fourth processing task for the fourth position based on part or all of the content in the third content except the fourth position.
[0038] Based on the above method, the electronic device can implement the training of the first model for uniformly processing multiple tasks.
[0039] In one possible design, the acquiring of training data includes: acquiring the third content, second indication information, and third indication information; wherein the second indication information is used to indicate the execution of the third processing task for the third position in the third content; the third indication information is used to indicate the execution of the fourth processing task for the fourth position in the third content; inputting the third content and the second indication information into the second model to determine the fifth content; or, inputting the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, inputting the second indication information into the second model to determine the fifth content; wherein the fifth content includes the content obtained by the second model performing the third processing task based on the input content; inputting the third content and the third indication information into the second model to determine the sixth content; or, inputting the content at the fourth position in the third content and the third indication information into the second model to determine the sixth content; or, inputting the third indication information into the second model to determine the sixth content; wherein the sixth content includes the content obtained by the second model performing the fourth processing task based on the input content; and determining the fourth content based on the fifth content and the sixth content.
[0040] Optionally, determining the fourth content based on the fifth content and the sixth content includes: combining the seventh content determined based on the fifth content and the eighth content determined based on the sixth content as the fourth content; wherein the seventh content is the fifth content or content obtained by modifying the fifth content in response to a received operation; and the eighth content is the sixth content or content obtained by modifying the sixth content in response to a received operation.
[0041] The above method for obtaining training data can be applied to the scenario of training a first model for processing multiple tasks. Through the above method, the electronic device can use the set second model to collect partial training data used as the result of task processing. Among them, by modifying the content obtained by processing the second model, namely the fifth content and the sixth content, according to user operations, the accuracy of the task processing result data in the training data can be improved through manual correction. It can further improve the accuracy of model training using training data, thereby obtaining a first model with better processing effect.
[0042] In one possible design, the model training of the set model according to the training data to obtain the first model includes: inputting the training data into the set model to determine a ninth content; wherein the ninth content includes the content obtained by the set model performing the third processing task on the third position in the third content; and adjusting the set model according to the third content and the ninth content to obtain the first model.
[0043] In the above method, the electronic device can adjust the set model based on the deviation between the task processing result with higher accuracy, i.e., the ninth content, and the task processing result obtained by the set model, i.e., the third content, so that the task processing result of the set model approaches the more accurate task processing result, and then train to obtain the first model with better processing effect.
[0044] In a possible design, the ninth content also includes content obtained by the set model performing the fourth processing task on the fourth position in the third content. In combination with the above method, a first model with better processing effect for processing multiple tasks can be trained.
[0045] Optionally, the set model is a model of a set network structure or a set generative AI model.
[0046] In one possible design, the first content includes at least one of the following: text, image, audio, video, and code.
[0047] The first content described in the above method may include at least one type of content. Therefore, based on the above method, single-task or multi-task processing for various types of content can be achieved, and better task processing results can be obtained.
[0048] In a possible design, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.
[0049] Among them, the semantic mask image can clearly and intuitively represent the location of different types of content in the image, which helps to improve the accuracy of image location recognition and task processing.
[0050] In one possible design, the first model is a generative artificial intelligence (AI) model. The generative AI model is capable of performing generative tasks. Applying the generative AI model to the above method can further expand the range of tasks the model can perform and improve the model's task execution capabilities.
[0051] In a second aspect, the present application provides a content processing system, which includes a terminal device and a server; the terminal device is used to: send first content, first location information and first task information to the server; wherein the first location information is used to indicate a first location in the first content, and the first task information is used to indicate the execution of a first processing task for the first location; the server is used to: receive the first content, the first location information and the first task information from the terminal device; input the first content, the first location information and the first task information into a first model to determine a second content, wherein the second content includes the content obtained by the first model performing the first processing task for the first location based on a first reference content, and the first reference content includes part or all of the content in the first content except the first location; send the second content to the terminal device; the terminal device is also used to: receive the second content from the server.
[0052] In this method, in the first content, the first reference content can serve as the context content of the first position. By inputting the first content, the first position information and the first task information into the first model, the server can enable the first model to refer to the context content of the first position when performing the first processing task indicated by the first task information for the first position indicated by the first position information, and thus obtain a more accurate processing result. Therefore, the above method helps to improve the accuracy of the results output when the server uses the first model to perform tasks, so that the output results are more likely to meet the task processing requirements, thereby improving the practicality of the first model. When the first model is an AI model, the above method can improve the effect of processing tasks through the AI model, thereby improving the practicality of the AI model. In addition, by sending the first content, the first position information and the first task information to the server, the terminal device can perform the corresponding processing tasks and obtain the corresponding task processing results based on this information through the server, and the terminal device can obtain the task processing results from the server. Therefore, the workload of task processing on the electronic device side can be reduced or avoided, thereby improving the overall processing efficiency of the electronic device side.
[0053] In a possible design, the first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content.
[0054] Optionally, if the first position is a position before the first sub-content, then in the content input into the first model, the first position information and the first task information are before the first sub-content; if the first position is a position after the first sub-content, then in the content input into the first model, the first position information and the first task information are after the first sub-content.
[0055] In a possible design, the first content also includes a second sub-content, the first position is a position between the first sub-content and the second sub-content, and the first reference content also includes the second sub-content.
[0056] Optionally, in the content input into the first model, the first position information and the first task information are located between the first sub-content and the second sub-content.
[0057] In one possible design, the first content includes a third sub-content, and the third sub-content is located at the first position; the second content specifically includes the content obtained by the first model performing the first processing task on the third sub-content based on the first reference content.
[0058] In one possible design, the first position information includes starting position information and ending position information; wherein, the starting position information is used to indicate the starting point of the first position, and the ending position information is used to indicate the ending point of the first position; the third sub-content is located between the starting point and the ending point in the first content.
[0059] In one possible design, in the content input into the first model, the third sub-content is between the starting position information and the ending position information; and / or, in the content input into the first model, the first task information is before the starting position information, or between the starting position information and the third sub-content, or between the third sub-content and the ending position information, or after the ending position information.
[0060] In one possible design, the terminal device is also used to: obtain the first content and first indication information before sending the first content, first location information and first task information to the server; wherein the first indication information is used to indicate the execution of the first processing task for the first location in the first content; and determine the first location information and the first task information based on the first content and the first indication information.
[0061] In one possible design, the terminal device is also used to: after obtaining the first content and the first indication information, and before determining the first location information and the first task information based on the first content and the first indication information, display the first content and the first indication information, and receive a first operation; wherein the first operation is used to indicate: converting the first indication information into the first location information and the first task information based on the first content.
[0062] In one possible design, when the terminal device obtains the first content and the first indication information, it is specifically used to: obtain and display the first content through a first application, and determine the first indication information based on a received second operation acting on the first content; wherein the second operation is used to indicate the first indication information; or, obtain, through the first application, the first content and the first indication information indicated by a received third operation; or, receive, through the first application, the first content and the first indication information sent by a second application; wherein the second application is used to generate the first content and the first indication information.
[0063] Optionally, the terminal device is also used to: before receiving the first content and the first indication information sent by the second application through the first application, obtain and display the first content through the second application, and determine the first indication information based on the received fourth operation acting on the first content.
[0064] Optionally, when the terminal device sends the first content, the first location information and the first task information to the server, it is specifically configured to: send the first content, the first location information and the first task information to the server through the first application.
[0065] In one possible design, the terminal device is further configured to: receive a third operation before sending the first content, the first location information, and the first task information to the server; wherein the third operation is used to instruct execution of a processing task based on the first content, the first location information, and the first task information. Optionally, the terminal device is further configured to: display the first content, the first location information, and the first task information before receiving the third operation.
[0066] In one possible design, the terminal device is also used to: send second location information and second task information to the server; wherein the second location information is used to indicate the second location in the first content, and the second task information is used to indicate the execution of a second processing task for the second location; the server is also used to: when the first content, the first location information and the first task information are input into the first model, the second location information and the second task information are input into the first model; the second content also includes: the content obtained by the first model executing the second processing task for the second location based on the second reference content; wherein the second reference content includes part or all of the content in the first content except the second location.
[0067] In one possible design, the server is further used to: before inputting the first content, the first location information and the first task information into the first model to determine the second content, obtain training data, and perform model training on the set model based on the training data to obtain the first model; wherein the training data includes third content, third location information, third task information and fourth content; wherein the third location information is used to indicate the third location in the third content, and the third task information is used to indicate the execution of a third processing task for the third location; the fourth content includes: content obtained by executing the third processing task for the third location based on part or all of the content in the third content except the third location.
[0068] In one possible design, when the server obtains training data, it is specifically used to: obtain the third content and second indication information; wherein, the second indication information is used to indicate the execution of the third processing task for the third position in the third content; determine the third position information and the third task information based on the second indication information; input the third content and the second indication information into the second model to determine the fifth content; or, input the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, when there is no content at the third position, input the second indication information into the second model to determine the fifth content; wherein, the fifth content includes the content obtained by the second model executing the third processing task based on the input content; and determine the fourth content based on the fifth content.
[0069] Optionally, when the server determines the fourth content based on the fifth content, it is specifically configured to: use the fifth content as the fourth content, or, in response to a received sixth operation, modify the fifth content to obtain the fourth content; wherein the sixth operation is used to instruct to modify the fifth content to the fourth content. The second model can be a set generative AI model.
[0070] In one possible design, the training data also includes fourth position information and fourth task information; wherein, the fourth position information is used to indicate the fourth position in the third content, and the fourth task information is used to indicate the execution of a fourth processing task for the fourth position; the fourth content also includes: content obtained by executing the fourth processing task for the fourth position based on part or all of the content in the third content except the fourth position.
[0071] In one possible design, when the server obtains training data, it is specifically used to: obtain the third content, second indication information and third indication information; wherein the second indication information is used to indicate the execution of the third processing task for the third position in the third content; the third indication information is used to indicate the execution of the fourth processing task for the fourth position in the third content; input the third content and the second indication information into the second model to determine the fifth content; or, input the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, input the second indication information into the second model to determine the fifth content; wherein the fifth content includes the content obtained by the second model performing the third processing task based on the input content; input the third content and the third indication information into the second model to determine the sixth content; or, input the content at the fourth position in the third content and the third indication information into the second model to determine the sixth content; or, input the third indication information into the second model to determine the sixth content; wherein the sixth content includes the content obtained by the second model performing the fourth processing task based on the input content; and determine the fourth content based on the fifth content and the sixth content.
[0072] Optionally, when the server determines the fourth content based on the fifth content and the sixth content, it is specifically used to: use the content obtained by combining the seventh content determined based on the fifth content and the eighth content determined based on the sixth content as the fourth content; wherein the seventh content is the fifth content or the content obtained by modifying the fifth content in response to the received operation; the eighth content is the sixth content or the content obtained by modifying the sixth content in response to the received operation.
[0073] In one possible design, when the server performs model training on the set model based on the training data to obtain the first model, it is specifically used to: input the training data into the set model to determine the ninth content; wherein, the ninth content includes the content obtained by the set model performing the third processing task on the third position in the third content; according to the third content and the ninth content, adjust the set model to obtain the first model.
[0074] In a possible design, the ninth content also includes content obtained by the set model performing the fourth processing task on the fourth position in the third content.
[0075] Optionally, the set model is a model of a set network structure or a set generative AI model.
[0076] In one possible design, the first content includes at least one of the following: text, image, audio, video, and code.
[0077] In a possible design, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.
[0078] In one possible design, the first model is a generative artificial intelligence (AI) model.
[0079] In a third aspect, the present application provides a model training method, which is applied to an electronic device, the method comprising: obtaining training data; wherein the training data comprises first content, first location information, first task information, and second content; wherein the first location information is used to indicate a first location in the first content, and the first task information is used to indicate the execution of a first processing task for the first location; the second content comprises: content obtained by executing the first processing task for the first location based on part or all of the content in the first content except the first location; and the first model is obtained by performing model training on a set model based on the training data. Optionally, the electronic device may be a server or a terminal device.
[0080] In this method, in the training data, part or all of the content of the first content except the first position can be used as the context content of the first position, and the second content includes the content obtained by performing the first processing task on the first position based on the context content. Therefore, the model trained based on the training data can have the ability to perform corresponding processing tasks on the position in combination with the context content of the position in the content. And performing task processing on the position in combination with the context content of the position can obtain more accurate processing results. Therefore, the above-mentioned model training method can make the accuracy of the output results of the trained first model when performing tasks higher, and it is easier to meet the task processing requirements, thereby improving the practicality of the trained first model. Therefore, the above-mentioned model training method can be used to train a model with better task processing effect.
[0081] In one possible design, the acquiring of training data includes: acquiring the first content and first indication information; wherein the first indication information is used to indicate the execution of the first processing task for the first position in the first content; determining the first position information and the first task information based on the first indication information; inputting the first content and the first indication information into the second model to determine the third content; or, inputting the content at the first position in the first content and the first indication information into the second model to determine the third content; or, when there is no content at the first position, inputting the first indication information into the second model to determine the third content; wherein the third content includes the content obtained by the second model executing the first processing task based on the input content; and determining the second content based on the third content.
[0082] Optionally, determining the second content based on the third content includes: using the third content as the second content, or, in response to a received first operation, modifying the third content to obtain the second content; wherein the first operation is used to instruct to modify the third content to the second content. The second model can be a set generative AI model.
[0083] The above method for obtaining training data can be applied to the scenario of training a first model for processing a single task. Through the above method, the electronic device can use the set second model to collect partial training data used as the task processing result. Among them, by modifying the content obtained by processing the second model, that is, the third content, according to user operations, the accuracy of the task processing result data in the training data can be improved through manual correction. It can further improve the accuracy of model training using training data, thereby obtaining a first model with better processing effect.
[0084] In one possible design, the training data also includes second position information and second task information; wherein, the second position information is used to indicate the second position in the first content, and the second task information is used to indicate the execution of a second processing task for the second position; the second content also includes: content obtained by executing the second processing task for the second position based on part or all of the content in the first content except the second position.
[0085] Based on the above method, the electronic device can implement the training of the first model for uniformly processing multiple tasks.
[0086] In one possible design, the acquiring of training data includes: acquiring the first content, first indication information, and second indication information; wherein the first indication information is used to indicate the execution of the first processing task for the first position in the first content; the second indication information is used to indicate the execution of the second processing task for the second position in the first content; inputting the first content and the first indication information into the second model to determine the third content; or, inputting the content at the first position in the first content and the first indication information into the second model to determine the third content; or, inputting the first indication information into the second model to determine the third content; wherein the third content includes the content obtained by the second model performing the first processing task based on the input content; inputting the first content and the second indication information into the second model to determine the fourth content; or, inputting the content at the second position in the first content and the second indication information into the second model to determine the fourth content; or, inputting the second indication information into the second model to determine the fourth content; wherein the fourth content includes the content obtained by the second model performing the second processing task based on the input content; and determining the second content based on the third content and the fourth content.
[0087] Optionally, determining the second content based on the third content and the fourth content includes: combining the fifth content determined based on the third content and the sixth content determined based on the fourth content as the second content; wherein the fifth content is the third content or content obtained by modifying the third content in response to a received operation; and the sixth content is the fourth content or content obtained by modifying the fourth content in response to a received operation.
[0088] The above method for obtaining training data can be applied to the scenario of training a first model for processing multiple tasks. Through the above method, the electronic device can use the set second model to collect partial training data used as the result of task processing. Among them, by modifying the content obtained by the second model processing, namely the third content and the fourth content, according to user operations, the accuracy of the task processing result data in the training data can be improved through manual correction. It can further improve the accuracy of model training using training data, thereby obtaining a first model with better processing effect.
[0089] In one possible design, the model training of the set model according to the training data to obtain the first model includes: inputting the training data into the set model to determine the seventh content; wherein the seventh content includes the content obtained by the set model performing the first processing task on the first position in the first content; and adjusting the set model according to the first content and the seventh content to obtain the first model.
[0090] In the above method, the electronic device can adjust the set model based on the deviation between the task processing result with higher accuracy, i.e., the seventh content, and the task processing result obtained by the set model, i.e., the second content, so that the task processing result of the set model approaches the more accurate task processing result, and then train to obtain the first model with better processing effect.
[0091] In one possible design, the seventh content also includes content obtained by the set model performing the second processing task on the second position in the first content. In combination with the above method, a first model with good processing effect for processing multiple tasks can be trained.
[0092] In one possible design, the first content includes at least one of the following: text, image, audio, video, and code.
[0093] The first content described in the above method may include at least one type of content. Therefore, based on the above method, a model that can better process single tasks or multi-tasks of various types of content can be trained.
[0094] In a possible design, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.
[0095] Among them, the semantic mask image can clearly and intuitively represent the location of different types of content in the image, which helps to improve the accuracy of image location recognition and task processing.
[0096] In one possible design, the set model is a model of a set network structure or a set generative AI model, and the first model is a generative artificial intelligence (AI) model. The generative AI model is capable of performing generative tasks, and by applying the generative AI model to the above method, the range of tasks that the trained model can perform can be further expanded, and the task execution capability of the trained model can be improved.
[0097] In a fourth aspect, the present application provides an electronic device, comprising a memory and one or more processors; wherein the memory is used to store computer program code, and the computer program code comprises computer instructions; when the computer instructions are executed by the one or more processors, the electronic device executes the method described in the above-mentioned first aspect or any possible design of the first aspect, or executes the method described in the above-mentioned third aspect or any possible design of the third aspect.
[0098] In a fifth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on an electronic device, the electronic device executes the method described in the above-mentioned first aspect or any possible design of the first aspect, or executes the method described in the above-mentioned third aspect or any possible design of the third aspect.
[0099] In a sixth aspect, the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on an electronic device, the electronic device executes the method described in the first aspect or any possible design of the first aspect, or executes the method described in the third aspect or any possible design of the third aspect.
[0100] In a seventh aspect, the present application provides a chip system comprising a processor and a memory, wherein the memory stores instructions; when the instructions are executed by the processor, the method described in the first aspect or any possible design of the first aspect is implemented, or the method described in the third aspect or any possible design of the third aspect is implemented. The chip system may be composed of a chip or may include a chip and other discrete components.
[0101] The beneficial effects of the second to seventh aspects mentioned above can be referred to the corresponding beneficial effects in the first, second or third aspects mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] FIG1 is a schematic diagram of a task processing method;
[0103] FIG2 is a schematic diagram of the hardware architecture of an electronic device provided in an embodiment of the present application;
[0104] FIG3 is a schematic diagram of a software architecture of an electronic device provided in an embodiment of the present application;
[0105] FIG4a is a schematic diagram of the architecture of a content processing system provided in an embodiment of the present application;
[0106] FIG4 b is a schematic diagram of the architecture of another content processing system provided in an embodiment of the present application;
[0107] FIG5 is a schematic diagram of a content processing method provided in an embodiment of the present application;
[0108] FIG6 is a schematic diagram of a method for editing content provided in an embodiment of the present application;
[0109] FIG7 is a schematic diagram of a content conversion method provided in an embodiment of the present application;
[0110] FIG8 is a schematic diagram of a content processing result provided by an embodiment of the present application;
[0111] FIG9 is a flow chart of a content processing method provided in an embodiment of the present application;
[0112] FIG10 is a schematic diagram of original content provided in an embodiment of the present application;
[0113] FIG11 is a schematic diagram of a content conversion method provided in an embodiment of the present application;
[0114] FIG12 is a schematic diagram of a content processing process provided in an embodiment of the present application;
[0115] FIG13 is a schematic diagram of a content conversion control method provided in an embodiment of the present application;
[0116] FIG14 is a schematic diagram of a content conversion control interface provided in an embodiment of the present application;
[0117] FIG15 is a schematic diagram of a content conversion control interface provided in an embodiment of the present application;
[0118] FIG16 is a schematic diagram of a content conversion control interface provided in an embodiment of the present application;
[0119] FIG17 is a schematic diagram of a content processing method provided in an embodiment of the present application;
[0120] FIG18 is a schematic diagram of a content processing method provided in an embodiment of the present application;
[0121] FIG19 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0122] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0123] In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.
[0124] To facilitate understanding, exemplary descriptions of concepts related to this application are provided for reference.
[0125] 1) Electronic devices, which may be devices with processing and computing capabilities, for example, devices with AI models (such as generative AI models).
[0126] In some embodiments of the present application, the electronic device may be a computing device such as a server. For example, the electronic device may be a cloud server.
[0127] In some embodiments of the present application, the electronic device may also be a portable device, such as a mobile phone, a tablet computer, a wearable device with wireless communication function (such as a watch, a bracelet, etc.), a vehicle-mounted terminal device, augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), smart home devices (such as smart TVs, smart speakers, etc.), smart robots, workshop equipment, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, flying equipment (such as smart robots, drones, airplanes), etc.
[0128] Among them, a wearable device is a portable device that a user can wear directly on the body or integrate into the user's clothes or accessories.
[0129] In some embodiments of the present application, the electronic device may also be a portable terminal device that also includes other functions such as a personal digital assistant and / or a music player. Or a portable terminal device with other operating systems. The portable terminal device may also be other portable terminal devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present application, the electronic device may not be a portable terminal device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0130] 2) Generative AI (artificial intelligence) models are deep learning-based AI models that simulate human creative thinking and generate logical and coherent content such as text, images, audio, video, and code. Generative AI models can receive inputs such as text, images, audio, video, and code and generate new content in any of these forms.
[0131] It should be understood that in the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, a and b, a and c, b and c, or a, b and c, where a, b, c can be single or multiple.
[0132] Currently, the task execution performance of some AI models (such as generative AI models and large language base models) is still unstable. When executing certain tasks (such as complex tasks or related tasks), they may not achieve good task processing results, resulting in the low practicality of generative AI models. In addition, some AI models perform well in processing single tasks with a single input sequence, but perform poorly in processing complex tasks that include multiple tasks with a single input sequence.
[0133] For example, as shown in the schematic diagram (a) in Figure 1, when the complex task set by the user for a piece of text includes multiple editing tasks that the user annotates on the piece of text (for example, the five tasks shown in the schematic diagram (a) in Figure 1), if the piece of text and the corresponding multiple editing tasks are directly input into the AI model as a sequence, the output result of the AI model is difficult to meet all the requirements of multiple editing tasks at the same time, and therefore the processing effect is poor.
[0134] Currently, a multi-task splitting solution can also be used, that is, a complex task containing multiple tasks is split into multiple individual tasks that can be solved in a single input, and then the individual tasks are input into the AI model for processing, and then the processing results of each task are obtained separately. For example, this solution can be used to split the complex task shown in the schematic diagram (a) in Figure 1 into 5 separate tasks, namely Task 1 to Task 5, as shown in the schematic diagram (b) in Figure 1. By inputting the 5 separate tasks into the AI model 5 times, the processing results of the 5 separate tasks can be obtained separately, and after combining them, the processing result of the entire complex task can be obtained.
[0135] In the above method, after the complex task is split into individual tasks, the AI model can only process the task based on the limited content corresponding to the individual tasks, and the accuracy of the processing may not be guaranteed, resulting in an inability to obtain a good processing effect. In some cases, multiple tasks may be related (such as the "ditto" editing task commonly used in text annotations), and the split processing will also affect the processing effect. In addition, the above method requires each task to be processed separately, so the overall processing efficiency is low.
[0136] In summary, the current use of AI models for task processing is poor, especially for multi-tasking, and the efficiency is low. Therefore, the practicality of AI models is low.
[0137] Based on the above problems, in order to improve the practicality of the AI model, the embodiments of the present application provide a content processing method, system and electronic device. This solution can use the AI model to process tasks more simply, efficiently and accurately, thereby improving the efficiency and processing effect of the AI model processing tasks and improving the practicality of the AI model.
[0138] The technical solutions provided in the embodiments of this application can be executed by any computing device with processing and computing capabilities, or by a system composed of multiple computing devices with processing, computing, and communication capabilities. The computing device can be an electronic device, etc. For an introduction to the performance of the electronic device, please refer to the description in the above conceptual description or the relevant description below.
[0139] The following description uses the application of the technical solution of this application in an electronic device or in a system including multiple electronic devices as an example. The implementation process of the application in other computing devices is similar and will not be repeated. Optionally, the electronic device may have an AI model.
[0140] 2 , the structure of an electronic device to which the method provided in an embodiment of the present application is applicable is introduced.
[0141] As shown in Figure 2, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a SIM card interface 195, etc.
[0142] The sensor module 180 may include a gyroscope sensor, an acceleration sensor, a proximity light sensor, a fingerprint sensor, a touch sensor, a temperature sensor, a pressure sensor, a distance sensor, a magnetic sensor, an ambient light sensor, an air pressure sensor, a bone conduction sensor, and the like.
[0143] It is understood that the electronic device 100 shown in FIG2 is merely an example and does not limit the electronic device, and the electronic device may have more or fewer components than shown in the figure, may combine two or more components, or may have a different component configuration. The various components shown in FIG2 may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.
[0144] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. Among them, the controller can be the nerve center and command center of the electronic device 100. The controller can generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution.
[0145] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0146] The execution of the content processing method provided in the embodiment of the present application can be controlled by the processor 110 or completed by calling other components, such as calling the processing program of the embodiment of the present application stored in the internal memory 121, or calling the processing program of the embodiment of the present application stored in a third-party device through the external memory interface 120, to control the wireless communication module 160 to communicate data with other devices, thereby improving the intelligence and convenience of the electronic device 100 and enhancing the user experience. The processor 110 can include different devices. For example, when a CPU and a GPU are integrated, the CPU and the GPU can cooperate to execute the content processing method provided in the embodiment of the present application. For example, part of the algorithm in the content processing method is executed by the CPU, and another part of the algorithm is executed by the GPU to obtain faster processing efficiency.
[0147] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1. The display screen 194 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces (GUIs). For example, the display screen 194 can display photos, videos, web pages, or files.
[0148] In the embodiment of the present application, the display screen 194 can be an integrated flexible display screen, or a spliced display screen consisting of two rigid screens and a flexible screen located between the two rigid screens.
[0149] Camera 193 (either a front-facing camera or a rear-facing camera, or one camera serving as both) is used to capture still images or videos. Typically, camera 193 includes a photosensitive element, such as a lens assembly and an image sensor. The lens assembly includes multiple lenses (convex or concave) that capture light signals reflected from the object to be photographed and transmit the captured light signals to the image sensor. The image sensor generates an original image of the object to be photographed based on the light signals.
[0150] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the code of the operating system, application program (such as the function corresponding to the solution of the present application, etc.). The data storage area can store data created during the use of the electronic device 100, etc.
[0151] The internal memory 121 may also store one or more computer programs corresponding to the algorithms of the present application. The one or more computer programs are stored in the internal memory 121 and configured to be executed by the one or more processors 110. The one or more computer programs include instructions that can be used to perform the various steps in the following embodiments.
[0152] In addition, the internal memory 121 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0153] Of course, the code of the algorithm of the embodiment of the present application can also be stored in an external memory. In this case, the processor 110 can run the code of the algorithm of the embodiment of the present application stored in the external memory through the external memory interface 120.
[0154] A touch sensor, also known as a "touch panel," can be provided on the display screen 194. The touch sensor and the display screen 194 form a touch display screen, also known as a "touch screen." The touch sensor is used to detect touch operations applied to or near the touch sensor. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor can also be provided on the surface of the electronic device 100, at a location different from that of the display screen 194.
[0155] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0156] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0157] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110. In an embodiment of the present application, the mobile communication module 150 can also be used to exchange information with other devices.
[0158] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0159] The wireless communication module 160 can provide wireless communication solutions applied to the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (WiFi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be transmitted from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2. In the embodiment of the present application, the wireless communication module 160 can be used to establish a connection with other electronic devices and exchange data. Or the wireless communication module 160 can be used to access an access point device, send control instructions to other electronic devices, or receive data sent from other electronic devices.
[0160] In addition, the electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor. For example, music playback, recording, etc. The electronic device 100 can receive input from the key 190 and generate key signal input related to the user settings and function control of the electronic device 100. The electronic device 100 can use the motor 191 to generate a vibration prompt (such as an incoming call vibration prompt). The indicator 192 in the electronic device 100 can be an indicator light, which can be used to indicate the charging status, power changes, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 in the electronic device 100 is used to connect the SIM card. The SIM card can be inserted into the SIM card interface 195 or pulled out from the SIM card interface 195 to achieve contact and separation with the electronic device 100.
[0161] It should be understood that in actual applications, the electronic device 100 may include more or fewer components than those shown in FIG2 , and the embodiments of the present application are not limited thereto. The illustrated electronic device 100 is merely an example, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0162] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. For example, as shown in Figure 3, the software architecture can be divided into four layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the runtime and system library, and the (Linux) kernel layer.
[0163] The application layer is the top layer of the operating system, including native applications of the operating system, such as camera, gallery, calendar, Bluetooth, music, video, information, etc., and may also include third-party applications. The application involved in the embodiment of the present application is referred to as application (APP), which is a software program that can realize one or more specific functions. Typically, multiple applications can be installed in an electronic device, such as a camera application, a mailbox application, etc. The applications mentioned below can be system applications that are installed on the electronic device when it leaves the factory, or they can be third-party applications that the user downloads from the Internet or obtains from other electronic devices while using the electronic device.
[0164] Of course, developers can write applications and install them into this layer. In one possible implementation, applications can be developed using the Java language by calling the application programming interface (API) provided by the application framework layer. Developers can use the application framework to interact with the underlying layer of the operating system (such as the kernel layer) and develop their own applications.
[0165] The application framework layer provides the application API and programming framework. It includes predefined functions and can include a window manager, content provider, view system, telephony manager, resource manager, and notification manager.
[0166] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the display (or screen), capture the display, etc.
[0167] Content providers are used to store and retrieve data and make it accessible to applications. The data may include files (such as documents, videos, images, audio), text, and other information.
[0168] The view system includes visual controls, such as those that display text, images, and documents. The view system is used to build applications. The interface within a display window can be composed of one or more views. For example, the interface for a text notification icon might include a view that displays text and a view that displays an image.
[0169] The phone manager provides communication functionality for electronic devices. The notification manager enables applications to display notifications in the status bar, which can be used to convey informational messages and automatically disappear after a short period of time without user interaction.
[0170] The runtime includes the core library and the virtual machine. The runtime is responsible for the scheduling and management of the system.
[0171] The system's core library consists of two parts: one containing the Java language's callable functions and the other the system's core library. The application layer and application framework layer run within a virtual machine. For example, in Java, the virtual machine executes Java files from the application and framework layers as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0172] The system library can include multiple functional modules. For example: surface manager, media library, 3D graphics processing library (for example: OpenGL ES), 2D graphics engine (for example: SGL), image processing library, etc. The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.564, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis and layer processing, etc. The 2D graphics engine is a drawing engine for 2D drawing.
[0173] The kernel layer provides the operating system's core system services, such as security, memory management, process management, the network protocol stack, and the driver model. These services are all implemented at the kernel layer. The kernel layer also serves as an abstraction layer between the hardware and software stacks. This layer contains many drivers related to electronic devices, including the display driver, the keyboard driver for input devices, the Flash driver for memory-based devices, the camera driver, the audio driver, the Bluetooth driver, and the Wi-Fi driver.
[0174] It should be understood that the functional services described above are only examples. In actual applications, electronic devices can also be divided into more or fewer functional services according to other factors, or the functions of each service can be divided in other ways, or the functional services can be not divided but work as a whole.
[0175] The solutions provided in the embodiments of the present application are described in detail below.
[0176] Figure 4a is a schematic diagram of the architecture of a content processing system provided in an embodiment of the present application. The content processing system can be deployed in an electronic device. As shown in Figure 4a, the content processing system may include: a content acquisition unit, a content editing unit, a content conversion unit, an AI model, and a display unit.
[0177] The electronic device may be a server or a terminal device.
[0178] Exemplarily, the server may be a cloud server.
[0179] Exemplarily, when the electronic device is a terminal device, the hardware architecture of the terminal device can be implemented using the architecture shown in FIG2 , and the software architecture of the terminal device can be implemented using the architecture shown in FIG3 .
[0180] For example, when the software architecture of the terminal device is implemented using the architecture shown in FIG3 , the functional units in the content processing system can be respectively deployed in the architecture layers such as the application layer or the application framework layer shown in FIG3 , and no specific restrictions are made in the embodiments of the present application.
[0181] The content acquisition unit can be used to obtain the original content to be processed. The format of the original content to be processed can be various media formats such as text, images, audio, video, code, etc. The content acquisition unit can also be used to send the original content to the content editing unit. The original content can be understood as the content before editing.
[0182] The content editing unit can be used to interact with the user, thereby performing task editing on the original content based on the user's operations. The user can perform editing operations on any one or more locations in the original content to indicate tasks to be performed on any one or more locations in the original content. Each location may or may not contain content. Absence of content can also be understood as empty content. Based on each edit operation performed by the user on each location in the original content, the content editing unit can add task indication information to each location, indicating the task corresponding to the edit operation, to thereby obtain edited content. In some embodiments of the present application, the edited content may include: the original content; and task indication information for each of one or more tasks. Each of the one or more tasks can specifically be a task (or operation) performed on a location in the original content, and the task indication information for each task is used to indicate the task. In the above method, the original content may include one or more sub-contents, each of which may be part of or all of the original content. Each location in the original content may include a sub-content, or may not include a sub-content (which can also be understood as the sub-content included at that location being empty). Different sub-contents of the original content can be identical, partially identical, or completely different. When different sub-contents of the original content are identical, the tasks corresponding to the different sub-contents are different; when different sub-contents of the original content are partially identical or completely different, the tasks corresponding to the different sub-contents can be the same or different. The content editing unit can also be used to send the edited content to the content conversion unit.
[0183] In some embodiments of the present application, the one or more tasks may be tasks instructed by one user, or tasks instructed by multiple users.
[0184] The content conversion unit can be used to convert the edited content into single-task prompt content or multi-task prompt content. Specifically, when the edited content includes one task, the content conversion unit can convert the edited content into single-task prompt content, and when the edited content includes multiple tasks, the content conversion unit can convert the edited content into multi-task prompt content. In some embodiments of the present application, the single-task prompt content may include: the original content, and the task prompt content corresponding to the above-mentioned one task. The multi-task prompt content may include: the original content, and the task prompt content corresponding to each of the multiple tasks mentioned above. Specifically, the task prompt content corresponding to each task may include: the sub-content at the position corresponding to the task, the position information of the position corresponding to the task, and the task indication information of the task. Specifically, the position information is used to indicate the position. The content conversion unit can also be used to send the single-task prompt content or the multi-task prompt content to the AI model.
[0185] The AI model can be used to generate a single-task processing result based on single-task prompt content, or to generate a multi-task processing result based on multi-task prompt content. The single-task processing result includes the content obtained after executing one of the aforementioned tasks on the original content. The multi-task processing result includes the content obtained after executing multiple tasks on the original content. The single-task prompt content or the multi-task prompt content can be the input content of the AI model, and the single-task processing result or the multi-task processing result can be the output content of the AI model. In some embodiments of the present application, the AI model can be a trained model that can recognize and process single-task prompt content or multi-task prompt content. Regarding the training process of the AI model, please refer to the relevant description in the following embodiments and will not be described in detail here. In other embodiments of the present application, when the AI model is a model that cannot directly recognize and process multi-task prompt content, the content conversion unit can convert the multi-task prompt content into multiple single-task prompt contents, and then input each single-task prompt content into the AI model for processing, and obtain the single-task processing results corresponding to each single-task prompt content output by the AI model, and then summarize the content obtained after executing the multiple tasks on the original content.
[0186] The display unit can be used to display original content or edited content or single-task prompt content or multi-task prompt content or single-task processing results or multi-task processing results. In some embodiments of the present application, the display unit can be used to display the original content after the content acquisition unit acquires the original content. Alternatively, it can be used to display the edited content after the content editing unit obtains the edited content. Alternatively, it can be used to display the single-task prompt content or multi-task prompt content after the content conversion unit converts the edited content into single-task prompt content or multi-task prompt content. Alternatively, it can be used to display the single-task processing result or multi-task processing result after the AI model outputs the single-task processing result or multi-task processing result.
[0187] Figure 4b is a schematic diagram of the architecture of another content processing system provided in an embodiment of the present application. As shown in Figure 4b, the content processing system may include a terminal device and a server. Optionally, the terminal device may include a content acquisition unit, a content editing unit, a content conversion unit, a display unit, and a communication unit. The server may include a communication unit and an AI model.
[0188] Among them, the content acquisition unit, content editing unit, content conversion unit and display unit in the terminal device can refer to the above description and will not be repeated here. The communication unit in the terminal device can be used to send the single-task prompt content or multi-task prompt content determined by the content conversion unit to the communication unit in the server, and can also be used to send the single-task processing result or multi-task processing result from the communication unit in the server to other units. The communication unit in the server can be used to send the received single-task prompt content or multi-task prompt content to the AI model, so that the AI model processes the single-task prompt content or multi-task prompt content. The communication unit in the server can also send the single-task processing result or multi-task processing result determined by the AI model to the communication module of the terminal device. The AI model in the server can refer to the above description and will not be repeated here.
[0189] The content processing system shown in FIG. 4b differs from the content processing system shown in FIG. 4a in that the content processing system in FIG. 4a is deployed on the same electronic device, while the content processing system in FIG. 4b is distributed and deployed on different electronic devices. Furthermore, the architecture and internal functional unit structures and functions of the content processing system shown in FIG. 4b can be implemented with reference to the content processing system shown in FIG. 4a and will not be described in detail in this embodiment.
[0190] It should be understood that the architecture of the content processing system described above is only an example. In actual applications, electronic devices can also be divided into more or fewer functional units according to other factors, or the functions of each unit can be divided in other ways, or the functional units can be not divided and work as a whole.
[0191] Based on the above description, taking the content processing method provided in the embodiment of the present application as an example of being applied to the same electronic device, and taking the scenario in which the electronic device uses an AI model to process multiple tasks as an example, a possible content processing method provided in the embodiment of the present application can be referred to Figure 5. Regarding the scenario in which the electronic device uses an AI model to process a single task, the method shown in Figure 5 can be implemented and will not be described in detail in the embodiment of the present application.
[0192] As shown in FIG5 , a possible content processing method may include:
[0193] S501: The electronic device obtains original content and displays the original content.
[0194] In some embodiments of the present application, the original content may be content in various media formats such as documents, text, images, audio, video, and codes.
[0195] The electronic device may obtain original content by, for example, obtaining original content input by the user, generating original content based on user operations, or obtaining original content sent by other devices. The embodiments of the present application do not impose any specific restrictions on the electronic device's method of obtaining original content.
[0196] Example 1: In one example, taking the original content in text format as an example, the original content may be a text segment as shown in the schematic diagram (a) in Figure 6. In this example, *, x, X, etc. are all used to represent text content.
[0197] Optionally, when the method provided in an embodiment of the present application is applied to the content processing system shown in Figure 4a, the step of obtaining the original content described in step S501 can be performed by the content acquisition unit shown in Figure 4a, and the step of displaying the original content described in step S501 can be performed by the display unit shown in Figure 4a.
[0198] S502: The electronic device generates edited content according to multiple editing operations performed by the user on the original content, and displays the edited content.
[0199] In some embodiments of the present application, each edit operation performed by the user is used to select any location in the original content and indicate a task to be performed at that location. The edited content may include the original content and the task indication information corresponding to each of the multiple edit operations. The task indication information corresponding to each edit operation indicates the task indicated by the edit operation, that is, the processing task to be performed for the location selected by the edit operation.
[0200] Optionally, the edited content may further include a content mark corresponding to each editing operation in the multiple editing operations, wherein the content mark corresponding to each editing operation is used to mark the position selected by the editing operation.
[0201] When a user selects a location in the original content where content exists, selecting that location effectively selects the sub-content at that location. Therefore, in this scenario, selecting a location is equivalent to selecting a sub-content. Therefore, selecting a sub-content in the following embodiments of this application is equivalent to selecting the location where the sub-content is located.
[0202] In some embodiments of the present application, the positions or sub-contents selected by different editing operations may be exactly the same (which may also be understood as completely overlapping or coincident), or may be partially the same (which may also be understood as partially overlapping or coincident), or may be completely different (which may also be understood as completely non-overlapping or non-coinciding).
[0203] Based on the above method, the original content may include multiple sub-contents.
[0204] In a specific implementation, a user's editing operation on the original content may include selecting any location in the original content and performing an editing operation on that location. This editing operation may be used to indicate a task to be performed for that location. Optionally, the sub-content at that location may be a subset of the original content. In this case, each user editing operation is equivalent to issuing a task.
[0205] In some embodiments of the present application, the original content, location information, and task instruction information in the edited content may exist independently of each other; or part of the content may be integrated into a whole, while the other part of the content exists independently; or all of the content may be integrated into a whole, and no specific limitation is made in the embodiments of the present application.
[0206] As an optional embodiment, each time a user performs an editing operation, the electronic device can respond to the editing operation by adding task indication information corresponding to the editing operation (which can be used to indicate the task corresponding to the editing operation) to the original content and display the content obtained after adding the task indication information. The user can perform multiple editing operations on the original content according to the above method. After the user completes the multiple editing operations, the electronic device can obtain the edited content and display the edited content.
[0207] Example 2. In one example, based on Example 1 above, when a user performs an editing operation on the original content shown in the schematic diagram (a) of Figure 6 , as shown in the schematic diagram (b) of Figure 6 , the user's editing operation on the original content may include: selecting segment 1 in the original content and adding editing suggestion 1, "Processing A," to segment 1, to indicate that Processing A should be performed on segment 1. The task corresponding to this editing operation is: performing Processing A on segment 1. The sub-content corresponding to this editing operation is segment 1, and the task instruction information may be editing suggestion 1 shown in the schematic diagram (b) of Figure 6 . After the user performs the above-mentioned editing operation on the original content shown in the schematic diagram (a) of Figure 6 , the electronic device may switch from displaying the content shown in the schematic diagram (a) of Figure 6 to displaying the content shown in the schematic diagram (b) of Figure 6 , thereby presenting the user's editing operation on the original content to the user. Based on the same method as above, the user can sequentially edit the original content multiple times, and the electronic device can obtain and display the edited content after the user performs multiple edits.
[0208] In one example, the edited content displayed by the electronic device may be the content shown in the schematic diagram (a) in Figure 7. Based on the content, it can be determined that the user has performed three editing operations. Among them, the sub-content corresponding to the first editing operation is fragment 1, the corresponding task instruction information is editing opinion 1, and the corresponding task is: perform A processing on fragment 1. The sub-content corresponding to the second editing operation is fragment 2, the corresponding task instruction information is editing opinion 2, and the corresponding task is: perform B processing on fragment 2. The sub-content corresponding to the third editing operation is fragment 3, the corresponding task instruction information is editing opinion 3, and the corresponding task is: perform C processing on fragment 3.
[0209] Optionally, when the method provided in an embodiment of the present application is applied to the content processing system shown in Figure 4a, the step of generating edited content based on multiple editing operations performed by the user on the original content described in step S502 can be performed by the content editing unit shown in Figure 4a, and the step of displaying the edited content described in step S502 can be performed by the display unit shown in Figure 4a.
[0210] S503: In response to the received conversion operation, the electronic device converts the edited content into multi-task prompt content and displays the multi-task prompt content.
[0211] The conversion operation is used to indicate that the edited content is converted into the multi-task prompt content. In some embodiments of the present application, the process of converting the edited content into the multi-task prompt content may also be referred to as editing.
[0212] In some embodiments of the present application, the multi-task prompt content may include: the original content, task indication information corresponding to each of the multiple editing operations on the original content, and location information of the location corresponding to each of the multiple editing operations. The location corresponding to each editing operation is the location selected by the editing operation, and the location information of the location corresponding to each editing operation is used to indicate the location.
[0213] In other embodiments of the present application, the multi-task prompt content may be content obtained by adding location information of each location to the edited content.
[0214] In the above method, any location, the task indication information corresponding to the location, and the location information of the location can constitute the task prompt content of the location. Therefore, the multi-task prompt content can include the task prompt content of each location in the multiple locations.
[0215] In a possible solution, the location information of each location may include starting location information and ending location information, wherein the starting location information is used to indicate the starting point of the location, and the ending location information may be used to indicate the ending point (i.e., the ending point) of the location.
[0216] Optionally, the starting position information and the ending position information may be the same or different.
[0217] Optionally, the starting position information and the ending position information can be displayed as information in various formats such as marks, icons, text, symbols, etc. The embodiments of the present application do not specifically limit the formats of the starting position information and the ending position information.
[0218] Optionally, in the multi-task prompt content, the order of arrangement between any position, the starting position information of the position, the ending position information of the position, and the task instruction information corresponding to the position can be arbitrary or pre-set. For example, in one possible embodiment, when a sub-content exists at the position, the sub-content can be located between the starting position information of the position and the ending position information of the position. The task instruction information corresponding to the position can be located between the starting position information of the position and the position, or can be located between the position and the ending position information of the position.
[0219] In another possible solution, the location information of each location may include marking information for marking the location. The marking information may be, for example, an underline, an outline, a shadow, an icon, etc. The embodiment of the present application does not specifically limit the format of the marking information.
[0220] Example 3. In one example, based on Example 2 above, the electronic device converts the edited content shown in the schematic diagram (a) of FIG. 7 to obtain the multitasking prompt content shown in the schematic diagram (b) of FIG. 7 . Specifically, in the edited content shown in the schematic diagram (a) of FIG. 7 , the record of each user editing operation includes the corresponding segment and editing suggestion. In the multitasking prompt content shown in the schematic diagram (b) of FIG. 7 , the record of each user editing operation includes the corresponding segment, the segment's starting position information and ending position information, and the editing suggestion. For example, taking the order and format of the segment, the segment's starting position information and ending position information, and the editing suggestion in the multitasking prompt content as "[segment's starting position information][segment's editing suggestion (i.e., the segment's task instruction information)][segment][segment's ending position information]" as an example, as shown in the schematic diagram (b) of FIG. 7 , the task prompt content for segment 1 includes: [Starting position 1][Processing A] xxxxxxxx [Ending position 1]. Where [Starting Position 1] is the starting position information of Segment 1, [Processing A] is the task instruction information of Segment 1, xxxxxxxx is the content of Segment 1, and [Ending Position 1] is the ending position information of Segment 1. Since Segments 2 and 3 overlap, and Segment 2 comes before Segment 3, the task prompt content of Segment 2 includes the starting position information and task instruction information of Segment 3's task prompt content, and the task prompt information of Segment 3 includes the ending position information of Segment 2's task prompt information. As shown in the schematic diagram (b) of Figure 7, the combined task prompt content of Segments 2 and 3 is displayed as: [Starting Position 2][Processing B]xxxxxxx[Starting Position 3][Processing C]XXXX[Ending Position 2]*******[Ending Position 3]. Where [Starting Position 2] is the starting position information of Segment 2, [Processing B] is the task instruction information of Segment 2, xxxxxxx and XXXX are the content of Segment 2, and [Ending Position 2] is the ending position information of Segment 2. These contents constitute the task prompt content of Segment 2. [Starting position 3] is the starting position information of segment 3, [C processing] is the task instruction information of segment 3, XXXX and ******* are the contents included in segment 3, and [Ending position 3] is the ending position information of segment 3. These contents constitute the task prompt content of segment 3.
[0221] In some embodiments of the present application, the location information of each location can be associated with the sub-content at that location. Alternatively, the location information of each location can include identification information of that location. For example, in the multi-task prompt content shown in the schematic diagram (b) in FIG7 above, the sequence number "1" in the starting location information and the ending location information of segment 1 can be used as the identification information of segment 1, and the sequence number "2" in the starting location information and the ending location information of segment 2 can be used as the identification information of segment 2.
[0222] In some embodiments of the present application, when the position selected by the user when editing the original content does not include valid content in the original content, the sub-content at the position can be considered to be empty, and the task prompt content corresponding to the position may also not include the sub-content, that is, the task prompt content corresponding to the position may only include the editing opinions and location information corresponding to the position. Alternatively, the sub-content at the position can be considered to be the original content, and the task prompt content corresponding to the position includes the editing opinions, original content, and location information of the original content corresponding to the position. Alternatively, the sub-content at the position can be considered to be the content of the set range of the position where the user is editing, and the task prompt content corresponding to the position includes the editing opinions, the content of the set range, and location information of the content of the set range.
[0223] Optionally, when the method provided in an embodiment of the present application is applied to the content processing system shown in Figure 4a, the step of converting the edited content into multi-task prompt content in response to the received first operation described in step S503 can be performed by the content conversion unit shown in Figure 4a, and the step of displaying the multi-task prompt content described in step S503 can be performed by the display unit shown in Figure 4a.
[0224] S504: The electronic device generates a multitasking processing result according to the AI model and the multitasking prompt content.
[0225] Optionally, the AI model can be a generative AI model.
[0226] In some embodiments of the present application, the aforementioned AI model may be a trained AI model capable of recognizing multi-task prompt content. The AI model training process may refer to the methods described in the embodiments below and will not be described in detail here. The electronic device may directly input the multi-task prompt content as a single input sequence into the trained AI model and obtain the multi-task processing results output by the AI model.
[0227] As a possible solution, when the AI model processes the multi-task prompt content, it can determine the corresponding multiple tasks based on the multiple task prompt contents in the multi-task prompt content, and process each task separately based on the task prompt content of each task in the multiple tasks. When processing each task, the AI model can determine the task to be performed based on the editorial opinions of the task, and can determine the location corresponding to the task based on the location information corresponding to the task, as well as the contextual content of the location (that is, the content in the original content other than the selected location), and then perform the task to be performed for the location in combination with the contextual content of the location. Based on this method, the AI model can also refer to more content for processing when performing single-task processing, thereby improving the accuracy of processing and improving the processing effect of the AI model.
[0228] Example 4. In one example, based on Example 3 above, taking the multi-task prompt content as shown in the schematic diagram (b) of Figure 7 as an example, after the electronic device inputs the multi-task prompt content shown in the schematic diagram (b) of Figure 7 into the AI model, the multi-task processing result output by the AI model may be the content shown in Figure 8, which includes the content obtained after performing the corresponding tasks on the fragments 1 to 3 shown in the schematic diagram (b) of Figure 7 respectively. Among them, the fragment 10 shown in Figure 8 is the fragment obtained after performing the task of "performing A processing on fragment 1" on the fragment 1 shown in the schematic diagram (b) of Figure 7. The fragment 20 shown in Figure 8 is the fragment obtained after performing the task of "performing B processing on fragment 2" on the fragment 2 shown in the schematic diagram (b) of Figure 7, and the fragment 30 shown in Figure 8 is the fragment obtained after performing the task of "performing C processing on fragment 3" on the fragment 3 shown in the schematic diagram (b) of Figure 7.
[0229] Optionally, when the method provided in the embodiment of the present application is applied to the content processing system shown in FIG. 4 a , step S504 may be performed by the AI model shown in FIG. 4 a .
[0230] In the above method, the electronic device converts the content containing multiple tasks to be performed (i.e., edited content) into multi-task prompt content based on a certain format, and can describe and prompt multiple tasks to be performed separately through the relevant content in the multi-task prompt content, so as to facilitate AI model recognition. The electronic device can directly input the multi-task prompt content containing multiple tasks as a single input sequence into the AI model for processing, and obtain the corresponding multi-task processing results. While ensuring the processing effect, it can realize the execution of multiple tasks in a single input sequence, that is, the effect of obtaining the output results of multiple tasks through one input, thereby improving processing efficiency.
[0231] The method described in the above embodiment is exemplarily described below with reference to specific examples.
[0232] Taking the original content described in the above embodiment as the text content in the original document as an example, as shown in FIG9 , the process of a content processing method provided in an embodiment of the present application may include:
[0233] Step 1: The electronic device obtains the original document.
[0234] The original document includes the original content.
[0235] Example 5: In one example, the original content included in the original document may be the text content shown in FIG. 10 .
[0236] Step 2: The electronic device determines the edited document based on the interactive editing performed by the user on the original document.
[0237] The interactive editing may include selecting multiple segments in the original document and generating corresponding editing suggestions for each selected segment. The editing suggestions are used to indicate tasks to be performed on the segments.
[0238] The edited document obtained by the user interactively editing the original document includes edited content.
[0239] A segment can be a subset of the text content in the original document, and multiple segments can overlap. Each segment can correspond to at least one edit suggestion, and each edit suggestion corresponds to a task.
[0240] Exemplarily, as shown in FIG9 , the plurality of segments may include segment 1 to segment n (n is a positive integer greater than 1), and the editing opinions corresponding to segments 1 to segment n are editing opinion 1 to editing opinion n, respectively.
[0241] Example 6: In one example, based on Example 5 above, the user interactively edits the text content shown in Figure 10 to obtain the edited content, as shown in the schematic diagram (a) of Figure 11. The edited content includes five editing suggestions for five segments in the original content.
[0242] Step 3: The electronic device generates a multi-tasking prompt text based on the edited content.
[0243] The multi-task prompt text may be content obtained by marking the start and end positions of the editing segments and corresponding editing opinions on the original content in the original document.
[0244] Each edited segment in the multi-task prompt text has a corresponding start symbol, editing suggestion, and end symbol. The start symbol serves as the starting position information for the corresponding segment, and the end symbol serves as the ending position information for the corresponding segment. Each segment and its corresponding start symbol, editing suggestion, and end symbol constitute the task prompt content corresponding to that segment.
[0245] For example, as shown in FIG9 , the start symbol of the edited segment 1 may be Start Symbol 1, the editing opinion of the segment 1 may be Editing Opinion 1, and the end symbol of the segment 1 may be End Symbol 1. The start symbol of the edited segment 2 may be Start Symbol 2, the editing opinion of the segment 2 may be Editing Opinion 2, and the end symbol of the segment 2 may be End Symbol 2. The same applies to other segments.
[0246] Example 7. In one example, based on Example 6 above, the multi-task prompt text obtained after the edited content shown in the schematic diagram (a) in Figure 11 is converted by the electronic device can be the multi-task prompt text shown in the schematic diagram (b) in Figure 11. Among them, in the edited content shown in the schematic diagram (a) in Figure 11, each editing record of the user includes a segment area and editing opinions. The electronic device can convert each editing record of the user into “[ / Prmpt <id>[Editor's note] Original clip[ / Prmpt <id>]” format, and then obtain the multi-task prompt text shown in the schematic diagram (b) in Figure 11. <id><Editing comments>] as the beginning, indicating the starting position of the document fragment to be edited and the corresponding editing comments, [ / Prmpt <id>] as the end, indicating the end position of the document fragment to be edited, and the original fragment in the middle includes the edited fragment selected by the user, which belongs to a subset of the original content in the original document. For example, in the multi-task prompt text shown in the schematic diagram (b) of Figure 11, one of the edited fragments may include "Just like what Confucius said: "Extend yourself to others." The task prompt content corresponding to this fragment may be "[ / Prmpt]" as shown in the schematic diagram (b) of Figure 11. <1> <The allusion is used incorrectly, please correct>] Just like what Confucius said: "Extend your own feelings to others"[ / Prmpt <1> ]", where the first symbol in the task prompt is " / Prmpt <1> " is the start symbol, which can be used as the starting position information of the edited segment. The second symbol " / Prmpt <1> " is the end mark, which can be used as the end position information of the edited segment. The same applies to other segments.
[0247] Step 4: The electronic device inputs the multitasking prompt text into the AI model to obtain a multitasking processing result.
[0248] In one possible solution, for an AI model that can directly recognize multi-task prompt text (such as the trained AI model described in the aforementioned embodiment), the electronic device does not need to perform any other processing and directly inputs the obtained multi-task prompt text into the AI model to obtain the multi-task processing result output by the AI model.
[0249] In another possible solution, for an AI model that cannot recognize multi-task prompt texts, the electronic device can first generate multiple single-task prompt texts based on the multi-task prompt text, and input each single-task prompt text into the AI model for processing to obtain the task processing result corresponding to each single-task prompt text. Then, by summarizing and integrating the task processing results corresponding to multiple single-task prompt texts, the multi-task processing results corresponding to the multi-task prompt text can be obtained. Among them, each single-task prompt text may include a task prompt content of a segment corresponding to an editing opinion. Optionally, the segment in the task prompt content (that is, the segment corresponding to the editing opinion) can also be replaced with the content of the paragraph where the segment is located or the original content.
[0250] In another possible solution, for an AI model that cannot recognize multi-tasking prompt text, the electronic device can explain how to handle the multi-tasking prompt text through prompts in the system settings or user context, thereby assisting the AI model to better handle the multi-tasking prompt text.
[0251] For example, as shown in FIG9 , the multi-task processing result output by the AI model after processing the multi-task prompt text may include the content obtained after performing corresponding processing on each segment, such as New Segment 1 to New Segment n shown in FIG9 . New Segment 1 to New Segment n are the contents obtained after the AI model performs corresponding task processing on segments 1 to n, respectively.
[0252] The above method achieves more efficient document processing. The electronic device converts the document so that the AI model can perform tasks corresponding to multiple editorial opinions in a single input (i.e., a single input sequence), resulting in higher processing efficiency. Furthermore, compared to methods that break down complex tasks into multiple executable simpler tasks, the above method has richer context, resulting in better processing results.
[0253] Taking the original content described in the above embodiment as an original image as an example, as shown in FIG12 , the process of a content processing method provided in an embodiment of the present application may include:
[0254] Step 1: The electronic device acquires the original image.
[0255] Exemplarily, the original image may be the original image shown in FIG12 .
[0256] Step 2: The electronic device determines an edited image based on the interactive editing performed by the user on the original image.
[0257] The interactive editing may include selecting multiple image regions in the original image and generating editing opinions for each selected image region.
[0258] The edited image obtained after the user interactively edits the original image includes edited content.
[0259] The image region can be a subset of the original image, and multiple image regions can overlap. Each image region can correspond to at least one editing suggestion, and each editing suggestion corresponds to a task.
[0260] For example, the image obtained after the user interactively edits the original image shown in FIG12 may be the edited image shown in FIG12 . For example, the area selected by the user in the edited image may include a background area and a portrait area, wherein the editing suggestion 1 corresponding to the background area may be: change the background. The editing suggestion 2 corresponding to the portrait area may be: change the hairstyle.
[0261] Step 3: The electronic device generates multi-task prompt content based on the edited image.
[0262] Exemplarily, as shown in FIG12 , the multi-task prompt content generated by the electronic device based on the edited image may include: the original image, a semantic mask image obtained by semantically segmenting the original image, and editing opinions, namely, editing opinions 1 and editing opinions 2. The semantic mask image includes mask area 1 and mask area 2. Mask area 1 is the position information of the background area, which can be used to indicate the position of the background area in the original image. Mask area 2 is the position information of the portrait area, which can be used to indicate the position of the portrait area in the original image. Editing opinion 1 is the editing opinion corresponding to the background area, and editing opinion 2 is the editing opinion corresponding to the portrait area.
[0263] Optionally, the editing opinions may be separate text content, or may be markup content added to the mask region of the semantic mask image. In the embodiments of the present application, no specific limitation is imposed on the format and setting method of the editing opinions.
[0264] Step 4: The electronic device inputs the multitasking prompt content into the AI model to obtain the multitasking processing result.
[0265] For example, the content obtained after the multi-task prompt content is processed by the AI model may be the multi-task processing result image shown in Figure 12. The multi-task processing result image includes the content obtained by processing the background area and the portrait area in the original image according to the corresponding editing suggestions.
[0266] In the above method, users can select different image regions during the editing process and assign different editing tasks to each region. By converting the image, the electronic device enables the AI model to process multiple editing suggestions for the image and generate the results of multiple image editing tasks using a single input, resulting in higher processing efficiency.
[0267] Based on the above description, an embodiment of the present application further provides a content conversion control method, which can be applied to the aforementioned embodiment, and specifically can be applied to the process of the electronic device performing steps S501 to S503 described in the aforementioned embodiment.
[0268] 13 , the content conversion control method provided in the embodiment of the present application may be any of the following methods:
[0269] Method 1: The electronic device executes the above steps S501 to S503 through a first application.
[0270] In this method, the first application may generate or obtain original content and display the original content. After displaying the original content, the first application may generate edited content in response to a user's operation to edit the original content and display the edited content. After displaying the edited content, the first application may generate multitasking prompt content based on the edited content in response to a received conversion operation and display the multitasking prompt content.
[0271] In the method, a first application may be installed in the electronic device. Optionally, the first application may be an application for generating original content.
[0272] In some embodiments of the present application, when displaying edited content, the first application may display a conversion control for triggering content conversion. In response to a conversion operation received on the conversion control, the first application may generate and display multitasking prompt content based on the edited content. The conversion operation is used to instruct the edited content to be converted into multitasking prompt content.
[0273] For example, in the case where the original content is text content in the original document described in the above embodiment, the first application can be a document application. As shown in FIG14 , when displaying the edited content, the document application can display a conversion control, and in response to a user clicking the conversion control, can convert the edited content into multi-tasking prompt content and display the multi-tasking prompt content.
[0274] Optionally, the first application may display the multitasking prompt content by: displaying the multitasking prompt content after the original content, or replacing the original content with the multitasking prompt content, or displaying the multitasking prompt content in a new interface or a newly created document interface.
[0275] Method 2: The electronic device executes steps S501 to S502 through a first application and executes step S503 through a second application.
[0276] The second application can be a system application or a third-party application. For example, the system application can be a system-type AI processing application such as a smart assistant. The third-party application can be a third-party AI processing application installed on the electronic device.
[0277] In this method, the first application may generate or obtain original content and display the original content. After displaying the original content, the first application may generate edited content in response to a user editing operation on the original content and display the edited content. After obtaining the edited content, the second application may convert the edited content into multitasking prompt content in response to a received conversion operation.
[0278] The second application may obtain the edited content in any of the following ways:
[0279] 1) After displaying the edited content, the first application may forward the edited content to the second application in response to a received forwarding operation.
[0280] For example, the forwarding operation may be an operation in which the user drags the edited content (or a file containing the edited content, such as a document) displayed by the first application to the application interface of the second application.
[0281] 2) The second application may obtain the edited content indicated by the user operation in response to the received user operation.
[0282] For example, a user can copy edited content (or a file containing edited content, such as a document) from a first application and paste the edited content into a second application, so that the second application obtains the edited content. In this example, the user operation received by the second application can be the user's operation of pasting the original content into the second application.
[0283] In some embodiments of the present application, after receiving the edited content, the second application can display the edited content. After converting the edited content into multi-tasking prompt content, the second application can display the multi-tasking prompt content. Optionally, when displaying the edited content, the second application can generate multi-tasking prompt content based on the edited content in response to the received conversion operation. The conversion operation is used to indicate that the original content is converted into multi-tasking prompt content. Exemplarily, the conversion operation can be an operation in which the user inputs a conversion indication, or can be an operation in which the user acts on a conversion control displayed by the second application. The conversion control can be used to trigger the conversion of the edited content into multi-tasking prompt content.
[0284] Exemplarily, taking the second application as an intelligent assistant application, and taking the original content as the text content in the original document described in the above embodiment as an example, the first application can be a document application. As shown in Figure 15, the document application can display the edited content, and the intelligent assistant application can display an application interface for interacting with the user. After displaying the edited content, the document application can send the document to the intelligent assistant application in response to the user dragging the document to which the edited content belongs to the application interface displayed by the intelligent assistant application. The intelligent assistant application can display the document in the user interface. And in response to the user's conversion operation (for example, the user enters the instruction "Optimize the document according to the annotations" in the application interface of the intelligent assistant application), the edited content in the document can be converted into multi-task prompt content, and the multi-task prompt content (for example, the document optimization draft shown in Figure 15) can be fed back.
[0285] Method 3: The electronic device executes the above step S501 through the first application, executes the above step S502 through the second application, and executes the above step S503 through the second application or the third application.
[0286] The first application may generate or obtain original content, and may also display original content. The second application may be a system application or a third-party application. The third application may be a system application or a third-party application.
[0287] In one possible method, a first application can generate or obtain original content and display the original content. After displaying the original content, the first application can forward the original content to a second application in response to a received forwarding operation. The specific forwarding method can refer to the method described above in Method 2 for forwarding edited content from the first application to the second application, and will not be described in detail here.
[0288] In another possible method, the second application can respond to the received user operation and obtain the original content indicated by the user operation. The specific acquisition method can refer to the method in which the second application obtains the edited content according to the user operation in the above method 2, which will not be described in detail here.
[0289] After obtaining the original content based on any of the above possible methods, the second application may display the original content, and may generate edited content in response to a user's operation of editing the original content, and display the edited content.
[0290] In one possible method, after displaying the edited content, the second application can generate multitasking prompt content based on the edited content in response to the received conversion operation, and display the multitasking prompt content. The specific implementation of this method can refer to the above method 1 and will not be described in detail here.
[0291] In another possible method, after the second application displays the edited content, the third application can, after obtaining the edited content, convert the edited content into multitasking prompt content in response to the received conversion operation. The specific implementation of this method can refer to the above method 2 and will not be described in detail here.
[0292] For example, taking the second application as a note-taking application, and the original content as the content in the original document described in the above embodiment, the first application can be a document application. After the note-taking application obtains and displays the original content, the user can edit the original content displayed by the note-taking application by handwriting, thereby obtaining the edited content shown in Figure 16.
[0293] It should be noted that the application described in the above method can also be replaced by a service process or a functional unit, etc.
[0294] Based on the above method, electronic devices support content editing and multi-tasking through a variety of interactive methods. In actual applications, appropriate methods can be selected for processing based on specific scenarios, so they are highly flexible and practical.
[0295] The following describes the training method of the AI model described in the embodiments of the present application.
[0296] A possible AI model training method provided in an embodiment of the present application may include:
[0297] Step 1: The electronic device obtains the training data set.
[0298] In this step, the training data set acquired by the electronic device may include single-task prompt content and corresponding single-task processing results, and / or multi-task prompt content and corresponding multi-task processing results. The format and specific contents of the single-task prompt content, single-task processing results, multi-task prompt content, and corresponding multi-task processing results can be found in the description of the previous embodiment and will not be repeated here.
[0299] After obtaining the single-task prompt content, the electronic device may input the single-task prompt content into a deep learning model for processing to obtain an output result corresponding to the single-task prompt content. After obtaining the multi-task prompt content, the electronic device may split the multi-task prompt content into multiple single-task prompt contents, and input each single-task prompt content into a deep learning model for processing to obtain an output result corresponding to each single-task prompt content. For each single-task prompt content, the electronic device may use the output result corresponding to the single-task prompt content as the single-task processing result corresponding to the single-task prompt content, or may use the modified output result as the single-task processing result corresponding to the single-task prompt content in response to a received operation. The operation is used to instruct the output result to be modified.
[0300] For multi-task prompt content, the electronic device can, after determining the single-task processing results corresponding to multiple single-task prompt contents in the multi-task prompt content, summarize and integrate the task processing results corresponding to the multiple single-task prompt contents to obtain the multi-task processing results corresponding to the multi-task prompt content, and then obtain the multi-task prompt content and the corresponding multi-task processing results for use as a training data set.
[0301] Optionally, the deep learning model can be an AI model with a set network structure or a set type. For example, the deep learning model can be an existing generative pre-trained transformer (GPT) model, a large language model, etc.
[0302] Step 2: The electronic device uses the training data set to train the deep learning model to obtain a trained AI model.
[0303] In one possible solution, the electronic device can input the multi-task prompt content in the training data set into a set model and use the set model to generate corresponding processing results. Based on the processing results of the set model and the single-task processing results and / or multi-task processing results in the training data, the set model can be subjected to reinforcement learning from human feedback (RLFH) to optimize the model parameters and processing effects of the set model, so that the set model learns to process tasks in combination with context and to process multiple tasks in a single input sequence. Optionally, the set model can be an AI model.
[0304] The model trained based on the above method can be used as a model that can recognize and process multi-task prompt content.
[0305] Of course, the electronic device can also use other training methods to train the set model, thereby obtaining an AI model that can combine context processing tasks and recognize and process multi-task prompt content. The specific model training method is not specifically limited in the embodiments of this application.
[0306] It should be noted that the implementation processes provided in the above embodiments are only examples of the method processes applicable to the embodiments of the present application. The execution order of each step can be adjusted accordingly according to actual needs, and other steps can be added or some steps can be reduced.
[0307] Based on the above embodiments and the same technical concept, an embodiment of the present application further provides a content processing method, as shown in FIG17 , which may include:
[0308] S1701: The electronic device obtains first content, first position information, and first task information; wherein the first position information is used to indicate a first position in the first content, and the first task information is used to instruct execution of a first processing task for the first position.
[0309] Optionally, the electronic device may be the electronic device shown in FIG4a , and the steps in the method may be executed by a content processing system deployed in the electronic device shown in FIG4a . The specific execution method may refer to the relevant description in the aforementioned embodiment and will not be repeated here.
[0310] The first content, the first location information, and the first task information are described in detail below.
[0311] In a first possible solution, the first content may include the first subcontent, the first position may be a position before or after the first subcontent, the first reference content may include the first subcontent, and the first reference content may be used as context content for the first position.
[0312] Exemplarily, the first content may be the original content described in the foregoing embodiments, the first position may be any position in the original content described in the foregoing embodiments, the first sub-content may be the content before or after the position, and the first task information may be the task indication information corresponding to the position (or the position information of the position).
[0313] Example 8. In one example, the first content may be the text content shown in FIG10 , the first sub-content may be "Be as resilient as building the Great Wall!" as shown in the schematic diagram (a) in FIG10 or FIG11 , and the first position may be the position before the first sub-content, that is, the position corresponding to the editorial opinion "Briefly introduce the project" as shown in the schematic diagram (a) in FIG11 .
[0314] Optionally, if the first position is a position before the first sub-content, then in the content input into the first model, the first position information and the first task information are before the first sub-content; if the first position is a position after the first sub-content, then in the content input into the first model, the first position information and the first task information are after the first sub-content.
[0315] Exemplarily, the content input into the first model may be the multi-task prompt content described in the above embodiment.
[0316] Example 9: In one example, based on Example 8 above, the content input to the first model may be the multi-task prompt text shown in the schematic diagram (b) of FIG11. The first position information includes two [ / Prmpt <0> ], the first task information is <Briefly introduce the project>. In the multi-task prompt text, the first position information and the first task information are located before the first sub-content.
[0317] In a second possible solution, based on the first possible solution described above, the first content may further include a second sub-content, the first position may be between the first sub-content and the second sub-content, and the first reference content may further include the second sub-content. Optionally, in the content input to the first model, the first position information and the first task information are located between the first sub-content and the second sub-content.
[0318] Example 10. In one example, the first content may be the text content shown in FIG10, the first sub-content may be "Hello everyone! Today I will tell you about the project we are about to start! We all know that this project is a big project," as shown in the schematic diagram (a) in FIG10 or FIG11, and the second sub-content may be "As resilient as building the Great Wall!" as shown in the schematic diagram (a) in FIG10 or FIG11, and the first position may be the position between the first sub-content and the second sub-content, that is, the position corresponding to the editorial opinion "Briefly introduce the project" shown in the schematic diagram (a) in FIG11. Correspondingly, the content input to the first model may be the multi-task prompt text shown in the schematic diagram (b) in FIG11. Among them, the first position information includes two [ / Prmpt <0> ], the first task information is <Briefly introduce the project>. In the multi-task prompt text, the first position information and the first task information are located between the first sub-content and the second sub-content.
[0319] In the third possible solution, the first content may include a third sub-content, and the third sub-content is located at the first position. Alternatively, based on the above-mentioned first possible solution or the above-mentioned second possible solution, the first content may also include a third sub-content, and the third sub-content is located at the first position (that is, the first position is the position where the third sub-content is located). In this solution, the second content specifically includes the content obtained by the first model performing the first processing task on the third sub-content based on the first reference content. In this solution, when there is substantial content in the first position of the first content (that is, the content is not empty), the content obtained by the first model performing the first processing task on the first position based on the first reference content is, that is, the content obtained by the first model performing the first processing task on the substantial content of the first position based on the first reference content.
[0320] Example 11. In one example, the first content may be the text content shown in FIG10 , and the third sub-content may be "Just like the lyrics of the song "We Are Different": "We are different, our dreams are different" shown in FIG10 ." The position of the third sub-content is the first position. The first sub-content may be part or all of the text content before the third sub-content shown in FIG10 , and the second sub-content may be part or all of the text content after the third sub-content shown in FIG10 .
[0321] In some embodiments of the present application, the first position information may include starting position information and ending position information; wherein the starting position information is used to indicate the starting point of the first position, and the ending position information is used to indicate the ending point of the first position; the third sub-content is located between the starting point and the ending point in the first content.
[0322] For example, in the above example 9 or example 10, the starting position information in the first position information may be the first [ / Prmpt] shown in the schematic diagram (b) in FIG11. <0> ](that is, before <Briefly introduce the project>[ / Prmpt <0> ]), the end position information in the first position information may be the second [ / Prmpt] shown in the schematic diagram (b) in FIG11. <0> ](i.e. the following section after <Briefly introduce the project> <0> ]).
[0323] Optionally, in the content input into the first model, the third sub-content may be located between the starting position information and the ending position information; and / or, in the content input into the first model, the first task information may be located before the starting position information, or between the starting position information and the third sub-content, or between the third sub-content and the ending position information, or after the ending position information.
[0324] For example, in Example 11 above, the content input to the first model is the multi-task prompt text shown in the diagram (b) of Figure 11. In the multi-task prompt text, the third sub-content (i.e., just like the lyrics of the song "We Are Different": "We are different, our dreams are different") is located at the starting position information (e.g. [ / Prmpt <4> ) and end position information (e.g. [ / Prmpt <4> The first task information (eg, <Change to an inspirational song>) is between the start position information and the third sub-content.
[0325] In some embodiments of the present application, the first content may include at least one of the following: text, image, audio, video, and code.
[0326] In one possible scenario, when the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image. For example, the first content (first image) may be the original image shown in FIG12 , the first position may be the position of the background shown in FIG12 , the first position information may be mask area 1 in the semantic mask image shown in FIG12 , and the first task information may be editing opinion 1 shown in FIG12 . For another example, the first content (first image) may be the original image shown in FIG12 , the first position may be the position of the person in the original image shown in FIG12 , the first position information may be mask area 2 in the semantic mask image shown in FIG12 , and the first task information may be editing opinion 2 shown in FIG12 .
[0327] The semantic mask image described in the embodiment of the present application refers to a specific form of image output generated by a semantic segmentation model for the original image, wherein each pixel in the semantic mask image is labeled as one of the predefined categories. The semantic segmentation model can generate a multi-label feature map (i.e., a semantic mask image) based on the original image, wherein the value of each pixel in the feature map represents the label of a different object category. Based on the semantic segmentation model, objects can be accurately classified and precise boundaries can be drawn at a fine pixel level.
[0328] The following describes in detail a method for obtaining the first content, the first location information, and the first task information.
[0329] In one possible solution, the electronic device may be a server, that is, the content display method provided in the embodiment of the present application may be applied to the server. Optionally, the server may be a cloud server. In this scenario, the electronic device may receive the first content, the first location information, and the first task information from the terminal device, and then obtain the first content, the first location information, and the first task information. Among them, the method for the terminal device to obtain the first content, the first location information, and the first task information can refer to the acquisition method when the electronic device is a terminal device in the previous or following text, which will not be described in detail here.
[0330] In another possible solution, the electronic device may be a terminal device, that is, the content display method provided in the embodiments of the present application may be applied to the terminal device. In this scenario, the electronic device may first obtain the first content and the first indication information, and then determine the first location information and the first task information based on the first content and the first indication information, thereby obtaining the first content, the first location information, and the first task information. The first indication information is used to instruct the execution of the first processing task for the first location in the first content.
[0331] Exemplarily, the first content may be the original content described in the foregoing embodiment, the first indication information may be the task indication information in the edited content described in the foregoing embodiment, and the first content and the first indication information may constitute the edited content described in the foregoing embodiment.
[0332] As an optional embodiment, after acquiring the first content and the first indication information, the electronic device may display the first content and the first indication information, and may receive a first operation before determining the first position information and the first task information based on the first content and the first indication information. The first operation is used to indicate: converting the first indication information into the first position information and the first task information based on the first content. The electronic device may determine the first position information and the first task information based on the first content and the first indication information in response to the received first operation. Optionally, after determining the first position information and the first task information based on the first content and the first indication information, the electronic device may display the first content, the first position information and the first task information so that the user can view the converted content to be input into the first model.
[0333] In some embodiments of the present application, the electronic device may obtain the first content and the first indication information in any of the following ways:
[0334] 1) Acquire and display first content through a first application, and determine first indication information according to a received second operation acting on the first content.
[0335] The first application is an application installed in the electronic device. The first application can be used to generate the first content, and can determine the first indication information according to the user's editing operation (ie, the second operation) on the first content.
[0336] Illustratively, the first application may be the first application described in method 1 of the content conversion control method described in the above embodiment, and the second operation may be the operation of editing the original content described in method 1 of the content conversion control method described in the above embodiment.
[0337] The specific implementation process of this method can be implemented by referring to method 1 of the content conversion control method described in the above embodiment, and will not be described in detail here.
[0338] 2) Acquire, through the first application, the first content and the first indication information indicated by the received third operation.
[0339] Illustratively, the first application may be the second application described in method 2 of the content conversion control method described in the above embodiment, and the second operation may be the forwarding operation or user operation described in method 2 of the content conversion control method described in the above embodiment.
[0340] The specific implementation process of this method can be implemented by referring to method 2 of the content conversion control method described in the above embodiment, and will not be described in detail here.
[0341] 3) Receiving, through the first application, first content and first indication information sent by the second application; wherein the second application is used to generate the first content and the first indication information.
[0342] In one example, the first application may be the second application described in method 2 of the content conversion control method described in the preceding embodiment, the second application may be the first application described in method 2 of the content conversion control method described in the preceding embodiment, and the second operation may be the forwarding operation described in method 2 of the content conversion control method described in the preceding embodiment. The specific implementation process of this method can be implemented with reference to method 2 of the content conversion control method described in the preceding embodiment and will not be described in detail here.
[0343] In another example, the first application may be the third application described in method 3 of the content conversion control method described in the preceding embodiment, and the second application may be the second application described in method 3 of the content conversion control method described in the preceding embodiment. The specific implementation process of this method can be implemented with reference to method 3 of the content conversion control method described in the preceding embodiment and will not be described in detail here.
[0344] S1702: The electronic device inputs the first content, the first location information and the first task information into the first model to determine the second content; wherein the second content includes: the content obtained by the first model performing the first processing task on the first location based on the first reference content; the first reference content includes part or all of the content in the first content except the first location.
[0345] In some embodiments of the present application, the first model is a generative artificial intelligence (AI) model.
[0346] Exemplarily, the first model may be the AI model described in the aforementioned embodiment.
[0347] The above method provides a content processing method in a single-task scenario (i.e., there is a set of location information and task information for the first content (the set of location information is the first location information and the first task information mentioned above)). In a multi-task scenario, there may be multiple sets of location information and task information for the first content. Each set of location information and task information includes location information for indicating a location in the first content and task information for indicating a processing task to be performed for the location. In a multi-task scenario, the electronic device can use the first model to uniformly process the first content and the corresponding multiple sets of location information and task information, and then obtain the task processing results corresponding to each set of location information and task information.
[0348] Based on this, in one possible scenario, the electronic device can also refer to the above-mentioned method for obtaining the first and second location information to obtain the second location information and second task information; wherein the second location information is used to indicate the second location within the first content, and the second task information is used to indicate the execution of the second processing task for the second location. When the electronic device inputs the first content, the first location information, and the first task information into the first model, it can also simultaneously input the second location information and the second task information into the first model. In other words, the electronic device can simultaneously input the first content, the first location information, the first task information, the second location information, and the second task information into the first model for processing. Based on this, the second content processed by the first model includes both the content obtained by the first model performing the first processing task for the first location based on the first reference content, and the content obtained by the first model performing the second processing task for the second location based on the second reference content; wherein the second reference content includes some or all of the content in the first content excluding the second location. Of course, the content input into the first model can include more sets of location information and task information, which can be implemented in accordance with the above-mentioned method and are not listed one by one in the embodiments of this application.
[0349] Among them, regarding the second position information, the second task information, the second reference information, etc., you can refer to the above description of the first position information, the first task information, the second reference information, etc., and will not be described in detail here.
[0350] The training process of the first model is described in detail below.
[0351] The training process of the first model may be performed before step S1702. In some embodiments of the present application, the training process of the first model may include the following steps 1 and 2.
[0352] Step 1: The electronic device obtains training data; wherein the training data includes third content, third position information, third task information and fourth content; wherein the third position information is used to indicate the third position in the third content, and the third task information is used to indicate the execution of the third processing task for the third position; the fourth content includes: content obtained by executing the third processing task for the third position based on part or all of the content in the third content except the third position.
[0353] In some embodiments of the present application, when the first model is used as a model in a single-task scenario, the training data may include original content (i.e., the third content mentioned above), a set of location information and task information (i.e., the third location information and third task information mentioned above) and corresponding task processing results (i.e., the fourth content mentioned above).
[0354] In some embodiments of the present application, when the first model is used as a model in a multi-task scenario, the training data may include original content, multiple sets of location information and task information and corresponding task processing results. Based on this, in a possible scenario, the training data includes the above-mentioned third content (as original content), third location information and third task information (as a set of location information and task information), and the fourth content. In addition, it may also include fourth location information and fourth task information (used as another set of location information and task information), and the fourth content also includes: content obtained by performing a fourth processing task on the fourth position based on part or all of the content in the third content except the fourth position; wherein the fourth location information is used to indicate the fourth position in the third content, and the fourth task information is used to indicate the execution of the fourth processing task for the fourth position. Of course, the training data may include more sets of location information and task information and corresponding task processing results, which can be implemented with reference to the above method, and will not be listed one by one in the embodiments of the present application.
[0355] In some embodiments of the present application, in the above-mentioned single-task scenario, the electronic device may obtain training data according to the methods described in the following steps A1 to A4.
[0356] A1: The electronic device obtains third content and second instruction information; wherein the second instruction information is used to instruct execution of a third processing task for a third position in the third content.
[0357] A2: The electronic device determines the third location information and the third task information according to the second instruction information.
[0358] A3: The electronic device inputs the third content and the second indication information into the second model to determine the fifth content; or, inputs the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, when there is no substantial content at the third position, inputs the second indication information into the second model to determine the fifth content; wherein the fifth content includes the content obtained by the second model performing the third processing task based on the input content.
[0359] Among them, the second model can be a generative AI model used in traditional solutions, such as an existing generative AI model, a large language model, etc.
[0360] Regarding the third content, second indication information, third location information and third task information mentioned above, reference may be made to the first content, first indication information, first location information and first task information described in the aforementioned embodiments, and they will not be described in detail here.
[0361] A4: The electronic device determines the fourth content according to the fifth content.
[0362] As an optional implementation, the electronic device may use the fifth content as the fourth content.
[0363] As another optional implementation, the electronic device can modify the fifth content in response to a received modification operation to obtain the fourth content; wherein the modification operation is used to indicate that the fifth content should be modified to the fourth content. Based on this method, the fourth content obtained by the second model processing can be manually modified to obtain the fifth content, which can ensure the accuracy of the task processing results. When the fifth content is used as training data, the accuracy of the training data can be guaranteed, thereby improving the model training effect.
[0364] In some embodiments of the present application, in the above-mentioned multi-tasking scenario, the electronic device may obtain training data according to the methods described in the following steps B1 to B4.
[0365] B1: The electronic device obtains third content, second indication information and third indication information; wherein the second indication information is used to instruct execution of a third processing task for a third position in the third content; and the third indication information is used to instruct execution of a fourth processing task for a fourth position in the third content.
[0366] B2: The electronic device inputs the third content and the second indication information into the second model to determine the fifth content; or, inputs the content at the third position in the third content and the second indication information into the second model to determine the fifth content; or, inputs the second indication information into the second model to determine the fifth content; wherein the fifth content includes the content obtained by the second model performing the third processing task based on the input content.
[0367] B3: The electronic device inputs the third content and the third indication information into the second model to determine the sixth content; or, inputs the content at the fourth position in the third content and the third indication information into the second model to determine the sixth content; or, inputs the third indication information into the second model to determine the sixth content; wherein the sixth content includes the content obtained by the second model performing the fourth processing task based on the input content.
[0368] The above steps B1 to B3 can be implemented with reference to the above steps A1 or A2, respectively, and will not be described in detail here.
[0369] B4: The electronic device determines the fourth content according to the fifth content and the sixth content.
[0370] Specifically, the electronic device may combine the seventh content determined based on the fifth content and the eighth content determined based on the sixth content as the fourth content. The seventh content is the fifth content or content modified from the fifth content in response to a received operation; the eighth content is the sixth content or content modified from the sixth content in response to a received operation. The specific implementation can refer to the above-mentioned method for determining the fourth content based on the fifth content and will not be described in detail here.
[0371] Step 2: The electronic device performs model training on the set model according to the training data to obtain a first model.
[0372] Specifically, the electronic device can input the training data obtained based on the above method into the set model to determine the ninth content. Specifically, in the above-mentioned single-task scenario, the ninth content includes the content obtained by the set model performing the third processing task on the third position in the third content. In the above-mentioned multi-task scenario, the ninth content also includes the content obtained by the set model performing the fourth processing task on the fourth position in the third content. After determining the ninth content, the electronic device can adjust the set model according to the third content and the ninth content to obtain the first model.
[0373] Exemplarily, the third content may be the single-task processing result / multi-task processing result in the training data described in the training method of the AI model in the aforementioned embodiment, and the ninth content may be the processing result of the set model described in the training method of the AI model in the aforementioned embodiment.
[0374] In the above method, the specific steps executed by the electronic device can also refer to the relevant introduction in the above embodiments, and will not be described in detail here.
[0375] Based on the above embodiments and the same technical concept, an embodiment of the present application further provides a content processing method, which can be applied to a content processing system including a server and a terminal device. As shown in FIG18 , the method may include:
[0376] S1801: The terminal device sends first content, first location information and first task information to the server; wherein the first location information is used to indicate a first location in the first content, and the first task information is used to indicate execution of a first processing task for the first location.
[0377] S1802: The server receives first content, first location information, and first task information from a terminal device.
[0378] S1803: The server inputs the first content, the first location information and the first task information into the first model to determine the second content, wherein the second content includes the content obtained by the first model performing the first processing task on the first location based on the first reference content, and the first reference content includes part or all of the content in the first content except the first location.
[0379] S1804: The server sends the second content to the terminal device.
[0380] S1805: The terminal device receives the second content from the server.
[0381] Optionally, the terminal device may be the terminal device shown in FIG4b, and the server may be the server shown in FIG4b. The steps in the method may be performed by a content processing system composed of the terminal device and the server shown in FIG4b. The specific execution method may refer to the relevant description in the aforementioned embodiment and will not be repeated here.
[0382] In the above method, regarding the specific steps executed by the terminal device or server, reference can be made to the execution method of the corresponding steps in the method shown in Figure 17, and no further details will be given here.
[0383] Based on the above embodiments and the same technical concept, an embodiment of the present application also provides a content processing system, which may include a server and a terminal device, the terminal device being used to: send first content, first location information and first task information to the server; wherein the first location information is used to indicate the first location in the first content, and the first task information is used to indicate the execution of a first processing task for the first location; the server being used to: receive the first content, first location information and first task information from the terminal device; input the first content, first location information and first task information into a first model to determine the second content, wherein the second content includes the content obtained by the first model executing the first processing task for the first location based on the first reference content, and the first reference content includes part or all of the content in the first content except the first location; send the second content to the terminal device; the terminal device is also used to: receive the second content from the server.
[0384] Among them, regarding the specific functions or execution methods of the terminal device and the server, please refer to the relevant descriptions in the aforementioned embodiments and will not be repeated here.
[0385] Based on the above embodiments and the same technical concept, an embodiment of the present application further provides an electronic device for implementing the content processing method provided in the embodiment of the present application for an electronic device, a server, or a terminal device. As shown in FIG19 , an electronic device 1900 may include: a memory 1901, one or more processors 1902, and one or more computer programs (not shown in the figure). The above-mentioned components may be coupled via one or more communication buses 1903. Optionally, the electronic device 1900 may further include a display screen 1904.
[0386] Among them, one or more computer programs (codes) are stored in the memory 1901, and one or more computer programs include computer instructions; one or more processors 1902 call the computer instructions stored in the memory 1901, so that the electronic device 1900 executes the content processing method applied to the electronic device or server or terminal device provided in the above-mentioned embodiment of the present application.
[0387] In a specific implementation, the memory 1901 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 1901 can store an operating system (hereinafter referred to as the system), such as an embedded operating system such as ANDROID, IOS, WINDOWS, or LINUX. The memory 1901 can be used to store the implementation program of the embodiment of the present application. The memory 1901 can also store a network communication program, which can be used to communicate with one or more additional devices, one or more user devices, and one or more network devices.
[0388] The one or more processors 1902 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.
[0389] Display screen 1904 is used to display application interface and other related user interfaces.
[0390] It should be noted that Figure 19 is only one implementation of the electronic device 1900 provided in an embodiment of the present application. In actual applications, the electronic device 1900 may also include more or fewer components. For details, please refer to the specific structure and description shown in Figure 3, which is not limited here.
[0391] Based on the above embodiments and the same technical concept, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer, the computer executes the method provided in the above embodiments and applied to an electronic device, server, or terminal device.
[0392] Based on the above embodiments and the same technical concept, an embodiment of the present application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer executes the method provided in the above embodiments and applied to an electronic device, server, or terminal device.
[0393] The methods provided in the embodiments of the present application may be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they may be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. A computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs), or semiconductor media (e.g., SSDs), etc.
[0394] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.< / id> < / id> < / id> < / id>
Claims
1. A content processing method, characterized in that: include: Acquire first content, first position information and first task information; wherein the first position information is used to indicate a first position in the first content, and the first task information is used to indicate execution of a first processing task for the first position; The first content, the first location information and the first task information are input into a first model to determine a second content; wherein the second content includes: the content obtained by the first model executing the first processing task for the first location based on the first reference content; the first reference content includes part or all of the content in the first content except the first location.
2. The method according to claim 1, characterized in that The first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content.
3. The method according to claim 2, characterized in that The first content also includes a second sub-content, the first position is a position between the first sub-content and the second sub-content, and the first reference content also includes the second sub-content.
4. The method according to any one of claims 1 to 3, characterized in that: The first content includes a third sub-content, and the third sub-content is located at the first position; the second content specifically includes content obtained by the first model performing the first processing task on the third sub-content based on the first reference content.
5. The method according to claim 4, characterized in that The first position information includes starting position information and ending position information; wherein the starting position information is used to indicate the starting point of the first position, and the ending position information is used to indicate the ending point of the first position; the third sub-content is located between the starting point and the ending point in the first content.
6. The method according to claim 5, characterized in that In the content input into the first model, the third sub-content is located between the start position information and the end position information; and / or In the content input to the first model, the first task information is located before the start position information, or between the start position information and the third sub-content, or between the third sub-content and the end position information, or after the end position information.
7. The method according to any one of claims 1 to 6, characterized in that: The method is applied to a server; the obtaining of the first content, the first location information and the first task information includes: The first content, the first location information, and the first task information are received from a terminal device.
8. The method according to any one of claims 1 to 6, characterized in that: The obtaining of the first content, the first location information and the first task information includes: Acquire the first content and first indication information; wherein the first indication information is used to instruct to perform the first processing task on the first position in the first content; The first location information and the first task information are determined according to the first content and the first indication information.
9. The method according to claim 8, characterized in that After acquiring the first content and the first indication information, and before determining the first location information and the first task information according to the first content and the first indication information, the method further includes: displaying the first content and the first indication information; A first operation is received; wherein the first operation is used to indicate: according to the first content, convert the first indication information into the first location information and the first task information.
10. The method according to claim 8 or 9, characterized in that The acquiring the first content and the first indication information includes: Acquire and display the first content through a first application, and determine the first indication information according to a received second operation acting on the first content; or acquiring, through the first application, the first content and the first indication information indicated by the received third operation; or The first content and the first indication information sent by the second application are received through the first application; wherein the second application is used to generate the first content and the first indication information.
11. The method according to any one of claims 8 to 10, characterized in that: The method is applied to a terminal device.
12. The method according to any one of claims 1 to 11, characterized in that: Before inputting the first content, the first position information and the first task information into the first model to determine the second content, the method further includes: acquiring second position information and second task information; wherein the second position information is used to indicate a second position in the first content, and the second task information is used to indicate executing a second processing task for the second position; When the first content, the first location information, and the first task information are input into the first model, the method further includes: inputting the second location information and the second task information into the first model; The second content also includes: content obtained by the first model executing the second processing task for the second position based on the second reference content; wherein the second reference content includes part or all of the content in the first content except the second position.
13. The method according to any one of claims 1 to 12, characterized in that: Before inputting the first content, the first location information, and the first task information into a first model to determine a second content, the method further includes: Acquire training data; wherein the training data includes third content, third position information, third task information and fourth content; wherein the third position information is used to indicate a third position in the third content, and the third task information is used to indicate the execution of a third processing task for the third position; the fourth content includes: content obtained by executing the third processing task for the third position based on part or all of the content in the third content except the third position; The first model is obtained by performing model training on the set model according to the training data.
14. The method according to claim 13, characterized in that The training data further includes fourth position information and fourth task information; wherein the fourth position information is used to indicate a fourth position in the third content, and the fourth task information is used to indicate the execution of a fourth processing task for the fourth position; The fourth content also includes: content obtained by executing the fourth processing task on the fourth position based on part or all of the content in the third content except the fourth position.
15. The method according to any one of claims 1 to 14, characterized in that: The first content includes at least one of the following: text, image, audio, video, and code.
16. The method according to any one of claims 1 to 15, characterized in that: When the first content is a first image, the first position information is a semantic mask image corresponding to the content at the first position in the first image.
17. A content processing system, characterized in that: include: Terminal devices and servers; The terminal device is used to: send the first content, the first position information and the first task information to the server; wherein the first position information is used to indicate the first position in the first content, and the first task information is used to indicate the execution of the first processing task for the first position; The server is used to: receive the first content, the first location information and the first task information from the terminal device; input the first content, the first location information and the first task information into a first model to determine a second content, wherein the second content includes content obtained by the first model performing the first processing task on the first location according to a first reference content, and the first reference content includes part or all of the content in the first content except the first location; and send the second content to the terminal device; The terminal device is further used for: receiving the second content from the server.
18. The system of claim 17, wherein: The first content includes a first sub-content, the first position is a position before or after the first sub-content, and the first reference content includes the first sub-content.
19. The system of claim 18, wherein: The first content also includes a second sub-content, the first position is a position between the first sub-content and the second sub-content, and the first reference content also includes the second sub-content.
20. The system according to any one of claims 17 to 19, characterized in that: The first content includes a third sub-content, and the third sub-content is located at the first position; the second content specifically includes content obtained by the first model performing the first processing task on the third sub-content based on the first reference content.
21. The system according to any one of claims 17 to 20, characterized in that: The terminal device is also used for: Before sending the first content, the first position information and the first task information to the server, obtaining the first content and first indication information; wherein the first indication information is used to indicate that the first processing task is to be performed for the first position in the first content; The first location information and the first task information are determined according to the first content and the first indication information.
22. The system of claim 21, wherein: The terminal device is also used for: After acquiring the first content and the first indication information, and before determining the first location information and the first task information according to the first content and the first indication information, displaying the first content and the first indication information, and receiving a first operation; The first operation is used to indicate: converting the first indication information into the first location information and the first task information according to the first content.
23. The system according to claim 21 or 22, characterized in that When the terminal device acquires the first content and the first indication information, it is specifically used to: Acquire and display the first content through a first application, and determine the first indication information according to a received second operation acting on the first content; or Acquire, through the first application, the first content and the first indication information indicated by the received third operation; or The first content and the first indication information sent by the second application are received through the first application; wherein the second application is used to generate the first content and the first indication information.
24. The system according to any one of claims 17 to 23, characterized in that: The terminal device is also used for: Before sending the first content, the first location information and the first task information to the server, receiving a third operation; The third operation is used to instruct to execute a processing task according to the first content, the first location information and the first task information.
25. The system according to any one of claims 17 to 24, characterized in that: The terminal device is further used to: send second position information and second task information to the server; wherein the second position information is used to indicate a second position in the first content, and the second task information is used to indicate execution of a second processing task for the second position; The server is further configured to: when the first content, the first location information and the first task information are input into the first model, input the second location information and the second task information into the first model; The second content also includes: content obtained by the first model executing the second processing task for the second position based on the second reference content; wherein the second reference content includes part or all of the content in the first content except the second position.
26. The system according to any one of claims 17 to 25, characterized in that: The server is also used to: Before inputting the first content, the first location information, and the first task information into the first model to determine the second content, obtaining training data, and performing model training on a set model according to the training data to obtain the first model; The training data includes a third content, a third position information, a third task information and a fourth content; wherein the third position information is used to indicate a third position in the third content, and the third task information is used to indicate the execution of a third processing task for the third position; the fourth content includes: content obtained by executing the third processing task for the third position based on part or all of the content in the third content except the third position.
27. The system of claim 26, wherein: The training data further includes fourth position information and fourth task information; wherein the fourth position information is used to indicate a fourth position in the third content, and the fourth task information is used to indicate the execution of a fourth processing task for the fourth position; The fourth content also includes: content obtained by executing the fourth processing task on the fourth position based on part or all of the content in the third content except the fourth position.
28. An electronic device, characterized in that: The electronic device includes a memory and one or more processors; The memory is used to store computer program codes, and the computer program codes include computer instructions; when the computer instructions are executed by the one or more processors, the electronic device executes the method according to any one of claims 1 to 16.
29. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 16.
30. A computer program product, characterized in that The computer program product comprises a computer program or instructions, and when the computer program or instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Method, electronic equipment and system for calling capabilities of other equipment
CN116171568A
Systems and methods for skip-based content detection
US11234031B1
Management of presentation content including interjecting live feeds into presentation content
US11263397B1
Dynamic presentation of searchable contextual actions and data
US20220309037A1
Methods and systems for adding annotations from a printed version of a document to a digital version of the document
US20230289515A1
Cited By
AI automatic script structured analysis video integrated generation method and system
CN122205200A