Data processing method and electronic equipment

By adjusting the display parameters of virtual objects, the problem of frame rate drop when computing power is strained was solved, ensuring output quality and improving user experience.

CN120909693APending Publication Date: 2025-11-07LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511073953.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

When virtual objects host meetings or give speeches, the strain on computing power can cause a drop in frame rate, affecting the output quality.

Method used

By adjusting the display parameters of virtual objects, such as frame rate, rendering resolution, and color fidelity, we can ensure that the frame rate of the output results meets the target frame rate and optimize computing power requirements.

Benefits of technology

It improves the output effect of virtual objects, avoids the frame rate drop problem caused by insufficient computing power, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909693A_ABST
    Figure CN120909693A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and electronic equipment, and the method comprises the steps: obtaining a target processing task which comprises to-be-output target multimedia data; processing the target processing task to obtain a first output result containing the target object; under the condition that the first frame rate of the first output result does not meet the target frame rate, display parameters of a target object are adjusted, a second output result is output, the second frame rate of the second output result meets the target frame rate, the target object is a virtual image capable of outputting target content, and the target content is content corresponding to the target multimedia data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a data processing method and electronic equipment. BACKGROUND

[0002] In daily life and work, it is more and more common to use virtual objects (such as digital people) to host meetings or speeches.

[0003] However, at present, it usually needs to consume high computing power to host meetings or speeches through virtual objects, and in the case of tight computing power, the frame rate of the virtual object will decrease, thereby affecting the output effect of the virtual object. SUMMARY

[0004] Therefore, the present application provides a data processing method and electronic equipment, and the specific scheme is as follows:

[0005] A data processing method applied to a first electronic device, comprising:

[0006] obtaining a target processing task, the target processing task comprising target multimedia data to be output;

[0007] processing the target processing task to obtain a first output result containing a target object;

[0008] in the case that a first frame rate of the first output result does not meet a target frame rate, adjusting a display parameter of the target object, and outputting a second output result, a second frame rate of the second output result meeting the target frame rate;

[0009] wherein the target object is a virtual image capable of outputting target content, and the target content is content corresponding to the target multimedia data.

[0010] Further, the processing the target processing task to obtain a first output result containing a target object comprises:

[0011] analyzing the target processing task through a target model to obtain a knowledge base of the target processing task and commentary information of the target object in the target processing task;

[0012] obtaining the first output result containing the target object based on the knowledge base and the commentary information.

[0013] Further, the analyzing the target processing task through the target model comprises at least one of the following:

[0014] in the case that the target multimedia data at least comprises first image information, identifying the content of the first image information through the target model to obtain the knowledge base and the commentary information.

[0015] In the case that the target multimedia data at least comprises text information, converting the text information into second image information by the target model, identifying the content of the second image information to obtain the knowledge base and the commentary information;

[0016] In the case that the target multimedia data at least comprises video information, identifying at least one video scene included in the video information by the target model, determining third image information corresponding to the at least one video scene in the video information, and identifying the content of the third image information to obtain the knowledge base and the commentary information.

[0017] Further, the analyzing the target processing task by the target model further comprises:

[0018] determining whether the target processing task contains critical information by the target model;

[0019] In the case that the target processing task contains critical information, processing the target processing task by the first electronic device;

[0020] In the case that the target processing task does not contain the critical information, processing the target processing task by the first electronic device and / or the second electronic device, the security level of the first electronic device is higher than that of the second electronic device, and the first electronic device and the second electronic device are connected by data.

[0021] Further, in the case that it is determined that the target processing task does not contain the critical information, the method further comprises:

[0022] In the case that it is determined that the first frame rate of the first electronic device outputting the first output result does not satisfy the target frame rate, switching to the second electronic device to process the target processing task;

[0023] Or,

[0024] In the case that it is determined that the first frame rate of the second electronic device outputting the first output result does not satisfy the target frame rate, switching to the first electronic device to process the target processing task.

[0025] Further, the adjusting the display parameter of the target object comprises at least one of the following:

[0026] adjusting the modulus information of the target object according to a preset frame rate threshold value and modulus information corresponding relationship table;

[0027] According to a first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and color restoration degree information corresponding relationship table, the color restoration degree information of the target object is adjusted.

[0028] According to a first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and control point quantity corresponding relationship table, the control point quantity of the target object is adjusted.

[0029] According to a first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and rendering resolution corresponding relationship table, the rendering resolution of the target object is adjusted.

[0030] Further, the adjustment of the display parameter of the target object further includes:

[0031] Reserve resources are determined, and the reserved resources are used to support the target object to output the second output result.

[0032] According to the current idle resources and the reserved resources, the display parameter of the target object is adjusted, so that the second frame rate of the second output result meets the target frame rate.

[0033] Further, it further includes:

[0034] Obtain the question information, the question information is the information to be answered obtained in the process that the target object outputs the target processing task processing result, and the question information includes audio information and image information.

[0035] The audio information and the image information are analyzed respectively to obtain audio analysis results and image analysis results.

[0036] Determine the question text data based on the audio analysis results and the image analysis results.

[0037] Query the knowledge base based on the question text data, and generate answer information according to the query result of the knowledge base, so that the target object outputs the answer information.

[0038] Further, the analysis of the audio information and the image information respectively to obtain the audio analysis results and the image analysis results includes:

[0039] The audio information is converted into first text data.

[0040] The question object included in the image information is analyzed by mouth shape analysis to determine second text data corresponding to the mouth shape of the question object, and the question object is an object outputting the question information.

[0041] The determining the question text data based on the audio analysis result and the image analysis result comprises:

[0042] performing semantic similarity analysis on the first text data and the second text data to obtain a similarity result;

[0043] in a case where the similarity result represents that the semantic similarity between the first text data and the second text data reaches a similarity threshold, taking any one of the first text data and the second text data as the question text data.

[0044] Further, the performing semantic similarity analysis on the first text data and the second text data to obtain a similarity result comprises:

[0045] in a case where the similarity result represents that the semantic similarity between the first text data and the second text data does not reach a similarity threshold, respectively determining a first natural language fluency value of the first text data and a second natural language fluency value of the second text data;

[0046] in a case where the first natural language fluency value is higher than the second natural language fluency value, taking the first text data as the question text data;

[0047] in a case where the first natural language fluency value is lower than the second natural language fluency value, taking the second text data as the question text data.

[0048] Further, the determining the question text data based on the audio analysis result and the image analysis result comprises:

[0049] determining a question language based on the audio analysis result and the image analysis result;

[0050] if it is determined that the question language is different from a target language, determining question language text data corresponding to the audio analysis result and the image analysis result, the target language being a language corresponding to the target object outputting the target processing task processing result;

[0051] converting the question language text data into text data in the target language, and determining the converted text data in the target language as the question text data.

[0052] An electronic device comprises:

[0053] The processor is configured to obtain a target processing task, the target processing task comprising target multimedia data to be output; process the target processing task to obtain a first output result comprising a target object; in a case where a first frame rate of the first output result does not satisfy a target frame rate, adjust a display parameter of the target object, and output a second output result, a second frame rate of the second output result satisfying the target frame rate; wherein the target object is a virtual image capable of outputting target content, and the target content is content corresponding to the target multimedia data.

[0054] The memory is configured to store programs required by the processor to perform the above processing process. BRIEF DESCRIPTION OF DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0056] Figure 1 A flowchart of a data processing method disclosed by an embodiment of the present application;

[0057] Figure 2 A flowchart of another data processing method disclosed by an embodiment of the present application;

[0058] Figure 3a A schematic diagram of commentary information disclosed by an embodiment of the present application;

[0059] Figure 3b Another schematic diagram of commentary information disclosed by an embodiment of the present application;

[0060] Figure 4 A flowchart of another data processing method disclosed by an embodiment of the present application;

[0061] Figure 5 A flowchart of another data processing method disclosed by an embodiment of the present application;

[0062] Figure 6 A flowchart of another data processing method disclosed by an embodiment of the present application;

[0063] Figure 7 A schematic diagram of the overall framework of the data processing method disclosed by an embodiment of the present application;

[0064] Figure 8 A schematic diagram of the structure of an electronic device disclosed by an embodiment of the present application. DETAILED DESCRIPTION

[0065] The embodiments of the present application are described below in conjunction with the accompanying drawings. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0066] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art can know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0067] The terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0068] The present application discloses a data processing method applied to a first electronic device, and a flowchart thereof is shown as Figure 1 As shown in the figure, the method comprises the following steps:

[0069] Step S11, obtaining a target processing task, the target processing task comprising target multimedia data to be output;

[0070] Step S12, processing the target processing task to obtain a first output result containing a target object;

[0071] Step S13, in the case that a first frame rate of the first output result does not meet a target frame rate, adjusting a display parameter of the target object, and outputting a second output result, a second frame rate of the second output result meeting the target frame rate, the target object being a virtual image capable of outputting target content, the target content being content corresponding to the target multimedia data.

[0072] In daily life and work, it is increasingly common to use virtual objects (such as digital humans) to host meetings or speeches.

[0073] However, at present, hosting a meeting or a speech through a virtual object usually requires high computing power, and in the case of tight computing power, the frame rate of the virtual object will decrease, thereby affecting the output effect of the virtual object.

[0074] Based on this, in the scheme, when the related content of the target multimedia data included in the target processing task needs to be output through the virtual image, the target processing task is processed, the frame rate of the output result is judged after the output result is obtained, and when the frame rate of the output result does not meet the target frame rate, the display parameters of the virtual image included in the output result are adjusted to make the frame rate of the adjusted output result meet the target frame rate, so as to ensure that the consumed computing power can match the target frame rate when the related content of the target multimedia data is output through the virtual image, thereby avoiding high computing power consumption and ensuring the output effect of the target object.

[0075] Specifically, if the target processing task is obtained, the target processing task includes target multimedia data to be output, such as text content, audio data, pictures, video data, etc., if the target object needs to output the related content of the target multimedia data included in the target processing task, the target processing task needs to be processed first to determine how the target object outputs and what content is output, that is, the target processing task is processed to obtain a first output result.

[0076] Among them, the target object is a virtual image capable of outputting target content, and the virtual image can be a digital person, that is, a virtual image with human appearance, behavior and interactive ability created by digital technology; the target content is the content corresponding to the target multimedia data, such as part or all of the content of the target multimedia data itself, or the content generated based on the target multimedia data, such as a commentary or a voice.

[0077] After obtaining the target processing task and determining that the target multimedia data to be output included in the target processing task needs to be output through the target object, when processing the target processing task, not only how the target object outputs and what content is output need to be determined, but also the specific virtual image of the target object needs to be determined to output the related content through the determined virtual image.

[0078] When the first output result is determined, the value of the first frame rate of the first output result is determined, and the value of the first frame rate is compared with the pre-set target frame rate to determine whether the first frame rate in the first output result meets the target frame rate. If the value of the first frame rate meets the target frame rate, it means that the current first output result meets the current computing power resources, and at this time, the target object in the first output result can directly output the related content of the target multimedia data according to the first output result.

[0079] If the value of the first frame rate does not meet the target frame rate, it indicates that the total demand computing power corresponding to the current first output result does not conform to the current computing power resource, at this time, the display parameter of the target object in the first output result can be adjusted, so as to adjust the total demand computing power of the target object in the output, so as to obtain the second output result.

[0080] The frame rate represents the number of image frames rendered per second, and its implementation mainly depends on the processing capability of the computing resource (CPU / GPU, etc.) for each frame of data, wherein the computing power demand of a single frame x frame rate = total computing power demand, therefore, the frame rate and the total computing power demand are in direct proportion, that is, when the computing power demand of a single frame is unchanged, the higher the frame rate, the higher the total computing power demand, and the lower the frame rate, the lower the total computing power demand.

[0081] The target frame rate can be a frame rate range, and not meeting the target frame rate can be not within the frame rate range, which can be greater than the maximum value of the frame rate range, or less than the minimum value of the frame rate range.

[0082] Adjusting the display parameter of the target object in the first output result can change the total demand computing power, if the value of the first frame rate is greater than the maximum value of the frame rate range, it indicates that the total demand computing power corresponding to the first output result is greater than the current computing power resource, therefore, at this time, the total demand computing power corresponding to the first output result needs to be reduced, in the embodiment, the adjustment of the demand computing power is realized by adjusting the display parameter of the target object, that is, the value of the display parameter of the target object is reduced, to obtain the second output result after adjusting the first output result, the second frame rate in the second output result meets the target frame rate.

[0083] If the value of the first frame rate is less than the minimum value of the frame rate range, it indicates that the total demand computing power corresponding to the first output result is less than the current computing power demand, therefore, at this time, the total demand computing power corresponding to the first output result can be increased, such as increasing the value of the display parameter of the target object, to obtain the second output result after adjusting the first output result, the second frame rate in the second output result meets the target frame rate.

[0084] It should be noted that whether the value of the display parameter of the target object is reduced or increased, through one adjustment of the value of the display parameter of the target object, the frame rate in the output result may not meet the target frame rate, therefore, the value of the display parameter of the target object can be adjusted multiple times, after each adjustment, it is determined whether the frame rate of the output result after one adjustment meets the target frame rate, until the frame rate of the output result after a certain adjustment meets the target frame rate, and then stop, and the output result after the adjustment at this time is determined as the second output result.

[0085] In a case where the first frame rate of the first output result does not meet the target frame rate, the display parameter of the target object in the first output result is adjusted, so that the display effect of the target object in the adjusted output result changes, such as the display definition of the target object. In addition, adjusting the display parameter of the target object can also change the virtual image of the target object, such as the image of the dress, long hair or short hair of the virtual image. Different virtual images require different computing power.

[0086] In addition, in a case where the frame rate of the first output result does not meet the target frame rate, the display parameter of the target object is adjusted, and other parameters such as the resolution, brightness and other data of the target multimedia data included in the first output result are further adjusted.

[0087] In the data processing method disclosed in the embodiment, in a case where the first frame rate of the first output result does not meet the target frame rate, it is determined that the total required computing power corresponding to the first output result does not match the current computing power resource. At this time, the total required computing power is adjusted by adjusting the display parameter of the target object in the first output result, so that the adjusted total required computing power can meet the current computing power resource, thereby avoiding the problem that the required computing power does not match the current computing power resource, resulting in excessive consumption of computing power resource and affecting the output effect of the target object.

[0088] The data processing method disclosed in the embodiment is applied to a first electronic device. After obtaining a target processing task including target multimedia data to be output, the target processing task is processed to obtain a first output result. The first output result includes a target object. If the first frame rate of the first output result does not meet the target frame rate, the display parameter of the target object in the first output result needs to be adjusted to obtain a second output result. The second frame rate of the second output result meets the target frame rate. The target object is a virtual image capable of outputting the corresponding content of the target multimedia data. When the target processing task is determined and it is determined that the related content of the target multimedia data included in the target processing task needs to be output by the virtual image, the target processing task is processed to obtain an output result. When the frame rate of the output result does not meet the target frame rate, the display parameter of the virtual image in the output result is adjusted, so that the frame rate of the adjusted output result can meet the target frame rate, thereby avoiding the problem that the adjusted output result consumes high computing power when outputting the related content by the virtual image, and further affecting the output effect of the virtual image. The user experience is improved.

[0089] The embodiment discloses a data processing method applied to a first electronic device, and a flowchart thereof is as shown in Figure 2 The embodiment discloses a data processing method applied to a first electronic device, and a flowchart thereof is as shown in

[0090] Step S21, obtaining a target processing task, the target processing task including target multimedia data to be output;

[0091] Step S22, parsing the target processing task by a target model to obtain a knowledge base of the target processing task and a side information of a target object in the target processing task;

[0092] Step S23, obtaining a first output result including the target object based on the knowledge base and the side information.

[0093] Step S24, in a case where a first frame rate of the first output result does not satisfy a target frame rate, adjusting a display parameter of the target object, and outputting a second output result, a second frame rate of the second output result satisfying the target frame rate, the target object being a virtual image capable of outputting target content, the target content being content corresponding to the target multimedia data.

[0094] After obtaining the target processing task, if content corresponding to target multimedia data included in the target processing task needs to be output by a target object, the target processing task needs to be processed to obtain a first output result including the target object, and it is determined whether a frame rate of the first output result satisfies a target frame rate, in a case where the frame rate of the first output result does not satisfy the target frame rate, a display parameter of the target object in the first output result is adjusted to change a required computing power when the content corresponding to the target multimedia data is output by the target object.

[0095] Specifically, processing the target processing task can specifically be processing the target processing task by a target model, the target model being used to process at least part of the target processing task, that is, the target model can process the target processing task to obtain a first output result including the target object, and in a case where a first frame rate of the first output result does not satisfy a target frame rate, a display parameter of the target object is adjusted to output a second output result, a second frame rate of the second output result satisfying the target frame rate.

[0096] Alternatively, the target model processes the target processing task to obtain a first output result including the target object, and the first output result is output, and whether the first frame rate of the first output result satisfies the target frame rate and the step of adjusting the display parameter of the target object are performed by other models or programs or software, not by the target model.

[0097] Specifically, processing the target processing task by the target model to obtain a first output result including the target object can be: parsing the target processing task by the target model to obtain a knowledge base of the target processing task and a side information of a target object in the target processing task; and obtaining the first output result including the target object based on the knowledge base and the side information.

[0098] The target model parses the target processing task after obtaining the target processing task, obtains all contents included in the target processing task, and stores the contents as a knowledge base of the target processing task, so that the target object can output the contents corresponding to the target multimedia data in the target processing task based on the contents of the knowledge base, or can reply based on the knowledge base when a question reply is needed.

[0099] In addition, the target model parses the target processing task, and can also obtain the voice-over information of the target object in the target processing task. The target object is a virtual image for outputting the contents corresponding to the target multimedia data included in the target processing task. When the target object outputs the contents corresponding to the target multimedia data, the contents output by the target object are the voice-over information. The voice-over information is related to the contents corresponding to the target multimedia data output by the target object.

[0100] The voice-over information can be related contents displayed on the display screen when the target object outputs the contents corresponding to the target multimedia data, such as text, pictures, or videos. At this time, the voice-over information is displayed on the display screen, as shown in FIG. 3, which includes a target object 31 and voice-over information 32. Figure 3a

[0101] The voice-over information can also be audio output by the target object when the target object outputs the contents corresponding to the target multimedia data. At this time, the voice-over information can be output only through audio and not displayed, or the display screen can also output synchronously while the audio is output, as shown in FIG. 4, which includes a target object 31 and voice-over information 32. The audio output by the target object is used to explain the displayed contents, or the display screen does not output contents and only the target object directly explains. Figure 3b

[0102] After the target model parses the target processing task, the knowledge base and the voice-over information of the target object are obtained, and based on the knowledge base and the voice-over information, it can be determined how the target object outputs the contents related to the target processing task, that is, the first output result is obtained.

[0103] For example, the target processing task is a file for introducing a product. After obtaining the file, the target model parses the file, determines all contents related to the product in the file, and forms a knowledge base. In addition, it can also determine the contents corresponding to the audio output by the virtual image when the virtual image explains the product introduced in the file, that is, the voice-over information. Based on the knowledge base and the voice-over information, it can be determined which contents need to be output through audio and which contents need to be output through the display screen when the virtual image explains the product, that is, the first output result is obtained.

[0104] ​​The target multimedia data included in the target processing task can be an image, text or video, that is, the target multimedia data can at least include first image information, can at least include text information, and can at least include video information.

[0105] In the case where the target multimedia data at least includes first image information, the content of the first image information is identified by the target model to obtain the knowledge base and the commentary information, that is, the image in the target multimedia data is identified by the target model to obtain the knowledge base and the commentary information.

[0106] In the case where the target multimedia data at least includes text information, the text information in the target multimedia data can be converted by the target model to obtain second image information, so that the target model can identify the second image information to obtain the knowledge base and the commentary information.

[0107] In the case where the target multimedia data at least includes video information, at least one video scene included in the video information is identified by the target model, third image information corresponding to the at least one video scene in the video information is determined, and the third image information is identified by the target model to obtain the knowledge base and the commentary information.

[0108] When the target multimedia data includes video information, the target model is used to determine how many video scenes are included in the video information. Different video scenes can determine different images, that is, different video scenes determine different third image information. For example, a video included in the target multimedia data includes three different video scenes, each of the three different video scenes determines at least one third image information, and only the determined third image information is identified to obtain the knowledge base and the commentary information. The third image information determined from each video scene is the key information in the video scene, and the frame image that is not determined as key information does not need to be identified and analyzed, so as to avoid the problem of excessive consumption of computing power caused by excessive data processing.

[0109] In addition, the target multimedia data included in the target processing task can only include text information, can only include image information, or can only include video information. Of course, it can include both text information and image information, or it can include both text information and video information, or it can include both image information and video information. It can also include text information, image information and video information. In this case, different identifications need to be performed to ensure the integrity of the content in the knowledge base and the commentary information.

[0110] The data processing method disclosed in the embodiment is applied to a first electronic device. When a target processing task is obtained, the target processing task is analyzed by a target model to obtain a knowledge base of the target processing task and side information of a target object in the target processing task, and a first output result containing the target object is obtained based on the knowledge base and the side information. In a case where a first frame rate of the first output result does not meet a target frame rate, a display parameter of the target object is adjusted, and a second output result is output. The target model is used to process the target processing task, so that the accuracy of the final output, i.e., the second output result, is ensured, and the computing power requirement of the second output result can meet the current computing power resource, thereby avoiding the problem that the computing power resource is excessively consumed and the output effect of the target object is affected.

[0111] The data processing method disclosed in the embodiment is applied to a first electronic device. When a target processing task is obtained, the target processing task is analyzed by a target model to obtain a knowledge base of the target processing task and side information of a target object in the target processing task, and a first output result containing the target object is obtained based on the knowledge base and the side information. In a case where a first frame rate of the first output result does not meet a target frame rate, a display parameter of the target object is adjusted, and a second output result is output. The target model is used to process the target processing task, so that the accuracy of the final output, i.e., the second output result, is ensured, and the computing power requirement of the second output result can meet the current computing power resource, thereby avoiding the problem that the computing power resource is excessively consumed and the output effect of the target object is affected. Figure 4

[0112] Step S41, a target processing task is obtained, and the target processing task includes target multimedia data to be output.

[0113] Step S42, whether the target processing task contains key information is determined by the target model.

[0114] Step S43, in a case where the target processing task contains the key information, the target processing task is processed by the first electronic device.

[0115] Step S44, in a case where the target processing task does not contain the key information, the target processing task is processed by the first electronic device and / or a second electronic device, the security level of the first electronic device is higher than that of the second electronic device, and the first electronic device and the second electronic device are data-connected.

[0116] After the target processing task is obtained, the target processing task is processed by the target model to obtain a first output result containing a target object, and whether a frame rate of the first output result meets a target frame rate is determined. In a case where the first output result does not meet the target frame rate, a display parameter of the target object in the first output result is adjusted to change the required computing power when the target multimedia data corresponding to the target object is output.

[0117] Before the target processing task is processed by the target model, whether the target processing task contains key information is determined by the target model. If the target processing task contains the key information, the target processing task is processed by the first electronic device. If the target processing task does not contain the key information, the target processing task is processed by the first electronic device and / or a second electronic device. The security level of the first electronic device is higher than that of the second electronic device, and the first electronic device and the second electronic device are data-connected.

[0118] ​The key information can be information containing privacy content, such as personal information (phone number, personal identification), address information, project information, company information, etc., or information in a keyword database, that is, a keyword database is set in advance, and the keyword database includes a plurality of keywords. When the target processing task contains at least one keyword in the keyword database, it can be determined that the target processing task contains key information.

[0119] The target model is used to analyze the target processing task to determine whether the target processing task contains key information, thereby ensuring the accuracy of the analysis of the target processing task.

[0120] The target model can analyze the title, theme, chapter title, etc. included in the target processing task to determine whether the target processing task contains key information, or directly analyze the target processing task to determine whether the content in the knowledge base contains key information after determining the knowledge base.

[0121] When the target model determines that the target processing task contains key information, a first electronic device with a higher security level can be selected to process the target processing task to ensure the security of the target processing task. When the target model determines that the target processing task does not contain key information, the security level of the electronic device executing the target processing task is not required. At this time, the first electronic device can be selected to process the target processing task, or the second electronic device can be selected to process the target processing task, or the first electronic device and the second electronic device can be used to process the target processing task.

[0122] When the target processing task does not contain key information, the electronic device executing the target processing task is no longer determined based on whether it contains key information, but is determined based on other information to determine which electronic device to use to execute the target processing task. For example, based on whether the idle computing resource of the first electronic device is greater than a target threshold value to determine whether the first electronic device can support the execution of the target processing task through the target model. If the idle computing resource of the first electronic device is greater than the target threshold value, the first electronic device can be selected to execute the target processing task. If the idle computing resource of the first electronic device is less than the target threshold value, the second electronic device can be selected to execute the target processing task. Further, if the idle computing resource of the first electronic device is less than the target threshold value, and the idle computing resource of the second electronic device is also less than the target threshold value, the first electronic device and the second electronic device can be selected to execute the target processing task together.

[0123] Further, in a case where it is determined that the first frame rate of the first electronic device outputting the first output result does not satisfy the target frame rate, switching to the second electronic device to process the target processing task; or in a case where it is determined that the first frame rate of the second electronic device outputting the first output result does not satisfy the target frame rate, switching to the first electronic device to process the target processing task.

[0124] If it is determined that the target processing task does not contain critical information, the first electronic device can be currently used to process the target processing task, that is, the target model is used to process the target processing task on the first electronic device, or the second electronic device can be currently used to process the target processing task, that is, the target model is used to process the target processing task on the second electronic device.

[0125] If the first electronic device is currently used to process the target processing task, a first output result containing a target object can be obtained, and whether the first frame rate of the first output result satisfies the target frame rate is determined. If it is determined that the first frame rate of the first output result does not satisfy the target frame rate, the first electronic device can be switched to the second electronic device, and the second electronic device performs subsequent processing tasks, that is, the first electronic device sends relevant information to the second electronic device, and the relevant information can include the target processing task and the first output result. The display parameters of the target object in the first output result are adjusted by the target model running on the second electronic device to output a second output result.

[0126] Alternatively, when the first electronic device is switched to the second electronic device, the second electronic device starts to execute the target processing task again, that is, the target model is used to process the target processing task on the second electronic device to obtain a first output result containing a target object. Whether the first frame rate of the first output result satisfies the target frame rate is determined. If the first frame rate of the first output result does not satisfy the target frame rate after switching to the second electronic device, the display parameters of the target object in the first output result currently obtained need to be adjusted.

[0127] Correspondingly, if the second electronic device is currently used to process the target processing task, a first output result containing a target object can be obtained, and whether the first frame rate of the first output result satisfies the target frame rate is determined. If it is determined that the first frame rate of the first output result does not satisfy the target frame rate, the second electronic device can be switched to the first electronic device, and the first electronic device executes subsequent processing tasks, or the first electronic device starts to execute the target processing task again.

[0128] It should be noted that the data connection between the first electronic device and the second electronic device can realize the transmission of data from the first electronic device to the second electronic device and the transmission of data from the second electronic device to the first electronic device, so as to ensure the switching of the electronic device executing the target processing task.

[0129] In addition, the first electronic device can be an entity electronic device, and the second electronic device can be a cloud server. The entity electronic device can process the target processing task in a manner of not connecting to a network, and thus has a higher security level than the cloud server. Alternatively, the second electronic device can also be an electronic device having a lower security level than the first electronic device, which is not limited herein.

[0130] If the second electronic device is a cloud server, the first frame rate of the first output result does not satisfy the target condition when the target processing task is executed by the target model using the second electronic device. This can be caused by insufficient idle computing power resources, or a poor network state.

[0131] Further, the target model can include a first model and a second model. The first model is configured to determine whether the target processing task contains key information, and determine to process the target processing task by the first electronic device if the target processing task contains the key information, or determine to process the target processing task by the second electronic device if the target processing task does not contain the key information. The second model is configured to process the target processing task on the corresponding electronic device to obtain a first output result containing a target object, adjust a display parameter of the target object if a first frame rate of the first output result does not satisfy a target frame rate, and output a second output result.

[0132] Alternatively, the target model includes only one large model. The large model can be used to determine whether the target processing task contains key information, determine the electronic device for executing the target processing task, and process the target processing task to finally output the second output result. The large model can be used to execute the complete target processing task in the data processing method disclosed in the embodiment, and can more accurately analyze and determine to ensure the accuracy of the finally output second output result.

[0133] The data processing method disclosed by the embodiment is applied to a first electronic device. After obtaining a target processing task, it is determined by a target model whether the target processing task contains key information. If the target processing task contains key information, the target processing task is processed by the first electronic device. If the target processing task does not contain key information, the target processing task is processed by the first electronic device and / or a second electronic device. The security level of the first electronic device is higher than that of the second electronic device, and the first electronic device and the second electronic device are connected in data. According to the scheme, after obtaining the target processing task, it is first determined whether the target processing task contains key information, so as to determine whether the target processing task is processed by the first electronic device or the second electronic device based on the determination. When the target processing task contains key information, the first electronic device with a higher security level is used to process the task, and when the target processing task does not contain key information, the electronic device used to process the task is not limited, thereby ensuring the security of the target processing task during execution and improving user experience.

[0134] The data processing method disclosed by the embodiment is applied to a first electronic device. After obtaining a target processing task, it is determined by a target model whether the target processing task contains key information. If the target processing task contains key information, the target processing task is processed by the first electronic device. If the target processing task does not contain key information, the target processing task is processed by the first electronic device and / or a second electronic device. The security level of the first electronic device is higher than that of the second electronic device, and the first electronic device and the second electronic device are connected in data. According to the scheme, after obtaining the target processing task, it is first determined whether the target processing task contains key information, so as to determine whether the target processing task is processed by the first electronic device or the second electronic device based on the determination. When the target processing task contains key information, the first electronic device with a higher security level is used to process the task, and when the target processing task does not contain key information, the electronic device used to process the task is not limited, thereby ensuring the security of the target processing task during execution and improving user experience. Figure 5

[0135] Step S51, obtaining a target processing task, the target processing task including target multimedia data to be output;

[0136] Step S52, processing the target processing task to obtain a first output result containing a target object;

[0137] Step S53, in a case where a first frame rate of the first output result does not meet a target frame rate, reserving a resource, the reserved resource being used to support the target object to output a second output result;

[0138] Step S54, adjusting a display parameter of the target object according to current idle resources and the reserved resource to obtain the second output result, a second frame rate of the second output result meeting the target frame rate, the target object being a virtual image capable of outputting target content, and the target content being content corresponding to the target multimedia data.

[0139] After obtaining the target processing task, if the content corresponding to the target multimedia data included in the target processing task needs to be output by the target object, the target processing task needs to be processed to obtain a first output result containing the target object. It is determined whether the frame rate of the first output result meets the target frame rate. In a case where the frame rate of the first output result does not meet the target frame rate, the display parameter of the target object in the first output result is adjusted to change the required computing power when the content corresponding to the target multimedia data is output by the target object, thereby ensuring the output effect of the target object.

[0140] ​The adjusting the display parameter of the target object in the first output result can specifically include: determining a reserved resource, the reserved resource being used to support the target object to output a second output result, and adjusting the display parameter of the target object according to the current idle resource and the reserved resource, so that a second frame rate of the second output result meets the target frame rate.

[0141] The size of the resource required for outputting the content corresponding to the target multimedia data by the target object is determined in advance or estimated, and the determined size of the resource is determined as the reserved resource, that is, this part of resource is reserved, so as to ensure that there is enough resource to support the target object to output the content corresponding to the target multimedia data, and avoid the situation that the resource is insufficient when the content corresponding to the target multimedia data is output by the target object.

[0142] The determining the reserved resource can specifically include: determining the reserved resource based on the narration information required when the first output result is determined. The determination of the first output result can include: parsing the target processing task to obtain a knowledge base of the target processing task and narration information of the target object in the target processing task, obtaining the first output result containing the target object based on the knowledge base and the narration information, and determining the reserved resource based on the narration information when the first frame rate of the first output result does not meet the target condition.

[0143] The narration information can be detailed narration information, or can be narration information only for key content. The narration information can generate different narration information based on the amount of idle resource, for example, when the idle resource is sufficient, detailed narration information can be generated, and when the idle resource is insufficient, narration information only for key content can be generated, so as to reduce the demand for resource.

[0144] When the first frame rate of the first output result does not meet the target condition, the reserved resource can be determined based on the narration information without adjusting the narration information, that is, whether the narration information is detailed narration information or narration information only for key content, the narration information is not adjusted, but the resource required for outputting the narration information by the target object is directly reserved;

[0145] Alternatively, the narration information can be adjusted based on the current idle resource to adjust how much resource is occupied by the narration information, so as to determine the size of the reserved resource. If the current idle resource is insufficient and the narration information is detailed narration information, the detailed narration information can be adjusted to be narration information only for key content, so as to reduce the occupation of resource by the narration information, and the size of the reserved resource is determined based on the adjusted narration information, so as to reserve the resource required for outputting the narration information by the target object.

[0146] Since the first frame rate in the first output result does not satisfy the target frame rate, it indicates that resources are not matched when the content corresponding to the target multimedia data is output through the target object in the first output result, i.e., the required resource when the content corresponding to the target multimedia data is output through the target object in the first output result does not match the current idle resource.

[0147] The required resource when the content corresponding to the target multimedia data is output through the target object in the first output result can be greater than the current required resource, in which case there is a resource shortage; or the required resource when the content corresponding to the target multimedia data is output through the target object in the first output result can be less than the current required resource, in which case the required resource is less, indicating that the first output result has a problem of a low first frame rate, and the first frame rate can be increased to ensure the output effect when the content corresponding to the target multimedia data is output through the target object.

[0148] Therefore, as long as it is determined that the first frame rate of the first output result does not satisfy the target frame rate and the display parameter of the target object in the first output result needs to be adjusted, the reserved resource and the current idle resource are determined, a part of the resource in the current idle resource is occupied as the reserved resource, and then the display parameter of the target object is adjusted based on the remaining resource area in the idle resource after the reserved resource is occupied, so that the second output result is obtained after the display parameter is adjusted, and the content corresponding to the target multimedia data is output through the target object in the second output result, which has sufficient resource support. At this time, the second frame rate of the second output result satisfies the target frame rate.

[0149] Further, when the display parameter of the target object in the first output result is adjusted, it can include at least one of the following:

[0150] adjusting the modulus information of the target object according to a preset frame rate threshold and modulus information corresponding relationship table;

[0151] adjusting the color restoration degree information of the target object according to the first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and color restoration degree corresponding relationship table;

[0152] adjusting the number of control points of the target object according to the first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and number of control points corresponding relationship table;

[0153] adjusting the rendering resolution of the target object according to the first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and rendering resolution corresponding relationship table.

[0154] That is, adjusting the display parameters of the target object can be achieved by adjusting the modulus information, color restoration degree information, control point quantity, rendering resolution, etc. of the target object. Adjusting the modulus information, color restoration degree information, control point quantity, rendering resolution, etc. of the target object can adapt to different computing resource, that is, by adjusting at least one of the above parameters of the target object, the size of the required resource when the target multimedia data is output by the target object can be adjusted.

[0155] The control point quantity can be the number of facial control points, the number of micro-expressions, the number of skeletal joints, etc. of the target object. By adjusting the control point quantity, the accuracy of the target object can be adjusted, thereby adjusting the required resource. Correspondingly, reducing the color restoration degree, reducing the rendering resolution, and reducing the modulus information can all achieve the purpose of reducing the required resource. Similarly, increasing the color restoration degree, increasing the rendering resolution, and increasing the modulus information can all achieve the purpose of increasing the required resource.

[0156] For the adjustment of the modulus information, a frame rate threshold and modulus information corresponding relationship table can be pre-set. When it is needed to adjust the frame rate to a certain frame rate range, the value of the modulus information corresponding to the frame rate range in the frame rate threshold and modulus information corresponding relationship table is found, and the modulus information of the target object in the first output result is adjusted to the determined value of the modulus information, that is, the adjustment of the frame rate is realized, and the adjustment of the required resource is realized.

[0157] For the adjustment of the color restoration degree, a frame rate threshold and color restoration degree corresponding relationship table can be set to adjust the color restoration degree once. Alternatively, the first frame rate and the target frame rate can be compared to adjust the color restoration degree multiple times, so that the adjusted frame rate meets the target frame rate. That is, when the first frame rate is greater than the target frame rate, it indicates that the total required computing resource corresponding to the first output result is greater than the current computing resource, and therefore, the color restoration degree information needs to be reduced. After reducing the color restoration degree information, it is continuously determined whether the current frame rate meets the target frame rate. If it still exceeds, the color restoration degree information needs to be continuously reduced until the adjusted frame rate meets the target frame rate. When the first frame rate is less than the target frame rate, the color restoration degree information can be increased. After increasing the color restoration degree information, it is continuously determined whether the current frame rate meets the target frame rate. If it still falls short, the color restoration degree information needs to be continuously increased until the adjusted frame rate meets the target frame rate.

[0158] For the adjustment of the number of control points, the frame rate threshold and the corresponding relationship table of the number of control points can be set to adjust to the right once, or the first comparison result of the first frame rate and the target frame rate can be adjusted multiple times, so that the adjusted frame rate meets the target frame rate. That is, when the first frame rate is greater than the target frame rate, the number of control points is reduced, and then it is continuously judged whether the current frame rate meets the target frame rate. If it is still greater, the number of control points needs to be continuously reduced until the adjusted frame rate meets the target frame rate. When the first frame rate is less than the target frame rate, the number of control points can be increased, and then it is continuously judged whether the current frame rate meets the target frame rate. If it is still less, the number of control points needs to be continuously increased until the adjusted frame rate meets the target frame rate.

[0159] For the adjustment of the rendering resolution, the frame rate threshold and the corresponding relationship table of the rendering resolution can be set to adjust to the right once, or the first comparison result of the first frame rate and the target frame rate can be adjusted multiple times, so that the adjusted frame rate meets the target frame rate. That is, when the first frame rate is greater than the target frame rate, the rendering resolution is reduced, and then it is continuously judged whether the current frame rate meets the target frame rate. If it is still greater, the rendering resolution needs to be continuously reduced until the adjusted frame rate meets the target frame rate. When the first frame rate is less than the target frame rate, the rendering resolution can be increased, and then it is continuously judged whether the current frame rate meets the target frame rate. If it is still less, the rendering resolution needs to be continuously increased until the adjusted frame rate meets the target frame rate.

[0160] The data processing method disclosed in this embodiment is applied to a first electronic device. After obtaining a target processing task, the target processing task is processed to obtain a first output result containing a target object. In the case where the first frame rate of the first output result does not meet the target frame rate, reserved resources are determined, and the display parameters of the target object are adjusted according to the current idle resources and the reserved resources to obtain a second output result. In this way, when the target object outputs the content corresponding to the target multimedia data included in the target processing task, sufficient resources can be provided for the execution, so that the problem of insufficient resources caused by excessive consumption of computing resources is avoided, and the output effect of the target object is further affected.

[0161] The data processing method disclosed in this embodiment is applied to a first electronic device. The flowchart is as shown in Figure 6

[0162] Step S61, obtaining a target processing task, the target processing task including target multimedia data to be output;

[0163] Step S62, parsing the target processing task to obtain the knowledge base of the target processing task and the voiceover information of the target object in the target processing task; ​

[0164] Step S63, obtaining a first output result containing the target object based on the knowledge base and the commentary information;

[0165] Step S64, in a case where a first frame rate of the first output result does not meet a target frame rate, adjusting a display parameter of the target object, outputting a second output result, a second frame rate of the second output result meeting the target frame rate, the target object being a virtual image capable of outputting target content, the target content being content corresponding to the target multimedia data;

[0166] Step S65, obtaining question information, the question information being information to be answered obtained in a process in which the target object outputs a target processing task processing result, the question information including audio information and image information;

[0167] Step S66, analyzing the audio information and the image information to obtain audio analysis result and image analysis result;

[0168] Step S67, determining question text data based on the audio analysis result and the image analysis result;

[0169] Step S68, querying the knowledge base based on the question text data, and generating answer information according to a query result of the knowledge base, so as to make the target object output the answer information.

[0170] After obtaining the target processing task, the target processing task needs to be processed (processing the target processing task can be performed by the target model) so as to obtain a second output result, and the content corresponding to the target multimedia data in the target processing task is output by the target object, wherein, in the process of obtaining the second output result, a knowledge base of the target processing task is obtained, and the knowledge base includes all information in the target processing task, and the information in the knowledge base is used as a basis for the target object to output the content corresponding to the target multimedia data.

[0171] Further, in the process of outputting the content corresponding to the target multimedia data by the target object, question information can be obtained, for example, a speech is made by a virtual image on a speech script, and a participant raises a question on the speech content in the process of making the speech, at this time, the system (a system based on the data processing method disclosed in the embodiment, or the target model) can obtain the question information, determine answer information through analysis of the question information, and output the answer information by the virtual image to complete answering the question of the participant.

[0172] Wherein, the participants can ask questions by voice or in the form of text. If the question is in the form of text, the system (the system based on the data processing method disclosed in this embodiment or the target model) can directly analyze the text question to determine the question text data. If the question is in the voice form, information collection is needed to convert the collected information to obtain the question text data.

[0173] For the case of asking questions in the voice form, the question information can include audio information and image information. That is, when determining that the question object outputs the question information, the question information needs to be collected. During the collection process, not only the audio of the voice output by the question object needs to be collected, but also the image when the question object outputs the voice needs to be collected. The collected audio and image are analyzed respectively to determine the question text data.

[0174] After determining the question text data, that is, the exact content that the question object needs to ask, the knowledge base can be queried to query the answer corresponding to the question text data from the knowledge base to generate the answer information and output it through the target object.

[0175] When outputting the answer information through the target object, the target object can be driven by the target algorithm to realize the output of the answer information. The target algorithm can be an audio2face algorithm (a technology for driving virtual character facial animation using audio signals. The core is to analyze audio signals through a deep learning model, extract key sound features, and map them to facial expressions and mouth movements). The audio2face algorithm can be connected to different virtual images bound by game development engines (such as unity and unreal engine, etc.). The algorithm and model are optimized respectively, and the information output by the audio2face algorithm is sent to the target object (the bound virtual image) to make the expression of the target object more vivid and the amplitude more full when outputting the content corresponding to the answer information and / or target multimedia data, so as to improve the user's viewing experience.

[0176] In this embodiment, since the question object asks questions in the voice form, there may be problems such as non-standard pronunciation or unclear sound when analyzing the received audio, which may cause the question object to ask the exact content of the question. Therefore, the collection of image information when the question object asks questions is increased, and the analysis of the image information is combined with the analysis of the audio information to finally determine the exact content of the question asked by the question object.

[0177] Specifically, the audio information can be converted into first text data, and the first text data is taken as an audio analysis result; a mouth shape of a questioning object included in the image information is analyzed to determine second text data corresponding to the mouth shape of the questioning object, and the second text data is determined as an image analysis result, and the questioning object is an object outputting the questioning information; semantic similarity analysis is performed on the first text data and the second text data to obtain a similarity result, and questioning text data is determined based on the similarity result.

[0178] For converting the audio information into the first text data, speech can be directly converted into text through automatic speech recognition (ASR). In addition, when the audio information is converted into the first text data, noise reduction processing, directional sound processing and the like can be added to ensure that the audio output by the questioning object can be clearly collected.

[0179] For converting the image information into the second text data, a mouth shape part of the questioning object in the image can be captured, and the mouth shape part is analyzed to determine the content of the questioning information determined through the image information.

[0180] The first text data and the second text data are subjected to semantic similarity analysis to determine whether the similarity reaches a specific similarity threshold to obtain a similarity result, so as to determine the questioning text data based on the similarity result.

[0181] If the similarity threshold is reached, it indicates that there is no big difference between the content of the first text data determined through the audio information and the content of the second text data determined through the image information, at this time, the first text data can be directly taken as the questioning text data, so as to query the knowledge base based on the first text data to generate answer information; or the second text data can be directly taken as the questioning text data, so as to query the knowledge base based on the second text data to generate answer information; in addition, the text data obtained by merging the first text data and the second text data can be taken as the questioning text data, and the knowledge base is queried based on this to generate answer information, so as to ensure the integrity of the content in the questioning text data and avoid omission.

[0182] The similarity threshold can be 0.6, 0.7, 0.75 or the like, which is only exemplarily described and not specifically limited.

[0183] Further, if the similarity result represents that the semantic similarity between the first text data and the second text data does not reach the similarity threshold, it indicates that there is a big difference between the content of the first text data determined through the audio information and the content of the second text data determined through the image information, at this time, the determination of the natural language fluency value can be increased.

[0184] That is, the natural language fluency values of the first text data and the second text data are determined respectively, and a comparison is made, and the question text data is determined based on the comparison result. Specifically, the first natural language fluency value of the first text data and the second natural language fluency value of the second text data are determined, and the first natural language fluency value and the second natural language fluency value are compared.

[0185] If the first natural language fluency value is greater than the second natural language fluency value, it indicates that the first text data is more in line with the language habits of natural language relative to the second text data, so that the conversion accuracy is higher when the audio data is converted into the first text data, and there is a high probability of conversion error when the image data is converted into the second text data. Therefore, at this time, the first text data can be directly used as the question text data, so as to select more accurate text data through semantic similarity and natural language fluency at the same time, thereby ensuring the accuracy of the question text data and also ensuring the accuracy of the answer information.

[0186] In addition, if the first natural language fluency value is less than the second natural language fluency value, it indicates that the second text data is more in line with the language habits of natural language relative to the first text data, so that the conversion accuracy is higher when the image data is converted into the second text data, and there is a high probability of conversion error when the audio data is converted into the first text data. Therefore, the second text data can be directly used as the question text data, so as to select more accurate text data through semantic similarity and natural language fluency at the same time, thereby ensuring the accuracy of the question text data and also ensuring the accuracy of the answer information.

[0187] The determination of semantic similarity can be realized by a word vector model Word2Vec model (the Word2Vec model can map words to vectors, capture semantic and grammatical relationships between words, and adjust word vector identifiers according to context), that is, the semantic similarity is calculated by the Word2Vec model, and it is determined according to the similarity calculation result whether to directly select one text data as the question text data or to perform natural language fluency judgment.

[0188] If the natural language fluency judgment needs to be performed, the NLP-Fluency model (an open source model focusing on evaluating natural language fluency) can be used to score the natural language fluency.

[0189] The data processing method disclosed in the embodiment comprises the following steps: when a target processing task is obtained, the target processing task is analyzed to obtain a knowledge base and side information, the second output result meeting a target frame rate is determined based on the knowledge base and the side information, the questioning information to be answered is obtained in the process of outputting the target processing task processing result by the target object, the questioning information comprises audio information and image information, the audio information and the image information are analyzed to obtain audio analysis result and image analysis result, the questioning text data is determined, the knowledge base is queried based on the questioning text data, the answering information is generated according to the query result of the knowledge base, and the answering information is output by the target object. In the scheme, if the questioning information to be answered is obtained in the process of outputting the target processing task processing result by the target object, the questioning text data is determined based on the analysis result of the audio information and the image information obtained when the questioning information is obtained, so as to obtain the answering information, which ensures the correctness of the determination of the questioning text information, improves the success rate of the correct answering of the answering information, and improves the user experience.

[0190] Further, the data processing method disclosed in the embodiment comprises the following steps:

[0191] The questioning language is determined based on the audio analysis result and the image analysis result. If the questioning language is determined to be different from the target language, the questioning language text data corresponding to the audio analysis result and the image analysis result is determined, the target language is the language corresponding to the output of the target processing task processing result by the target object, the questioning language text data is converted into the text data in the target language, and the converted text data in the target language is determined as the questioning text data.

[0192] In the process of speech and questioning, there may be different languages, for example, the language used for speech is Chinese, and the questioning language of the questioning object is English.

[0193] For this case, in the embodiment, the questioning language can be determined first. The questioning language can be determined based on only the audio analysis result, or can be determined based on only the image analysis result, or can be determined by simultaneously referring to the language obtained by analyzing the audio analysis result and the language obtained by analyzing the image analysis result. If the language obtained by analyzing the audio analysis result is the same as the language obtained by analyzing the image analysis result, the language obtained based on the audio analysis result or the image analysis result is determined as the questioning language.

[0194] After the question language is determined, it is determined whether the question language is the same as a target language, the target language being a language used by the target object to output content corresponding to the target multimedia data. If the question language is the same as the target language, no language switching is needed, and the question text data is determined based on the audio analysis result and the image analysis result. If the question language is different from the target language, the question language needs to be converted into the target language so as to unify the languages, and on this basis, the answer information corresponding to the question information is determined, and then the answer information is output by the target object, and the target language is still used when the target object outputs the answer information.

[0195] Specifically, if the question language is different from the target language, the question text data corresponding to the audio analysis result and the image analysis result is determined, that is, the question language is analyzed first to obtain the text data corresponding to the question language, for example, if the question language is English, English text data is determined. The determination of the text data corresponding to the question language can be specifically: converting the audio information of the question language into first text data of the question language, performing lip analysis on the question object included in the image information according to the question language to determine second text data of the question language corresponding to the lip shape of the question object, performing semantic similarity analysis on the first text data and the second text data of the question language to obtain a similarity result, and obtaining the text data corresponding to the question language based on the similarity result.

[0196] After the text data corresponding to the question language is obtained, the text data corresponding to the question language is converted into text data corresponding to the target language according to the conversion rule between the question language and the target language, so that the text data corresponding to the target language can be directly determined as the question text data, and then the knowledge base is queried to generate answer information corresponding to the question text data, the language corresponding to the answer information being the target language.

[0197] After the question information is obtained, the comparison between the language corresponding to the question information and the target language is performed, so that the target object can output the answer information corresponding to the question information according to the language required by the system.

[0198] Specifically, the data processing method disclosed in the embodiment has the overall framework as shown in FIG. 1. Figure 7As shown, it includes: an operation level, a function level and a model level, wherein for the operation level: first, document (i.e. target processing task) uploading and parsing, which is parsed by a large model (such as a large language LLM model), can store the document and generate a knowledge base through parsing; then generate the aside information, this step corresponds to the document skill at the operation level, generate the aside information based on the document through the large model, and the aside information can also be adjusted; then, generate the digital person (i.e. target object), in the generation process, the image of the digital person can be selected from the digital person image library, and the tone of the digital person can be selected from the digital person tone library, this process can be performed by a model, such as: the model generates the digital person image and tone based on the uploaded data, loaded jewelry and clothing, and photo modeling; then, document broadcasting, i.e. outputting the content corresponding to the target multimedia data in the target processing task through the target object, which can convert text into speech through a TTS algorithm model, and drive the digital person's mouth shape, facial expression, etc. to realize outputting the corresponding content through the digital person; there may be document interaction in the output process, i.e. obtaining question information, the digital person queries the database based on the question information, generates and outputs answer information, this process involves ASR speech recognition algorithm, TTS algorithm model and Audio2face algorithm model, the obtained speech is converted into text through the ASR speech recognition algorithm, AI interaction is performed through the large model, i.e. the large model analyzes and answers the obtained question information to obtain the answer information; in the process of AI interaction by the large model, the chat data of the interaction can also be stored; in addition, when the digital person outputs the related content, not only the question information can be proposed in the form of voice, but also the question information can be input on the display interface of the digital person outputting the related content, and the question information and the answer information determined based on the question information are displayed on the display interface.

[0199] The embodiment discloses an electronic device, a structural schematic diagram of which is shown as Figure 8 As shown, it includes:

[0200] The processor 81 and the memory 82.

[0201] The processor 81 is configured to obtain a target processing task, the target processing task including target multimedia data to be output; process the target processing task to obtain a first output result including a target object; in a case where a first frame rate of the first output result does not satisfy a target frame rate, adjust a display parameter of the target object, and output a second output result, a second frame rate of the second output result satisfying the target frame rate; wherein the target object is a virtual image capable of outputting target content, and the target content is content corresponding to the target multimedia data.

[0202] The memory 82 is configured to store programs required by the processor to perform the above processing process.

[0203] Further, the processor is configured to:

[0204] resolve the target processing task by the target model to obtain a knowledge base of the target processing task and a side information of a target object in the target processing task, and obtain a first output result containing the target object based on the knowledge base and the side information.

[0205] Further, the processor resolves the target processing task by the target model, including at least one of:

[0206] In a case where the target multimedia data at least includes first image information, the processor identifies a content of the first image information by the target model to obtain the knowledge base and the side information; in a case where the target multimedia data at least includes text information, the processor converts the text information into second image information by the target model, and identifies a content of the second image information to obtain the knowledge base and the side information; in a case where the target multimedia data at least includes video information, the processor identifies at least one video scene included in the video information by the target model, determines third image information corresponding to the at least one video scene in the video information, and identifies a content of the third image information to obtain the knowledge base and the side information.

[0207] Further, the processor is further configured to:

[0208] determine whether the target processing task contains critical information by the target model; in a case where the target processing task contains the critical information, process the target processing task by the first electronic device; in a case where the target processing task does not contain the critical information, process the target processing task by the first electronic device and / or the second electronic device, the security level of the first electronic device is higher than that of the second electronic device, and the first electronic device and the second electronic device are data-connected.

[0209] Further, in a case where it is determined that the target processing task does not contain the critical information, the processor is further configured to:

[0210] in a case where it is determined that a first frame rate at which the first electronic device outputs the first output result does not meet a target frame rate, switch to the second electronic device to process the target processing task; or in a case where it is determined that a first frame rate at which the second electronic device outputs the first output result does not meet the target frame rate, switch to the first electronic device to process the target processing task.

[0211] Further, the processor is configured to adjust a display parameter of the target object, including at least one of:

[0212] According to the preset frame rate threshold and the modulus information corresponding relationship table, the modulus information of the target object is adjusted; according to the first comparison result of the first frame rate and the target frame rate, or according to the preset frame rate threshold and the color restoration degree information corresponding relationship table, the color restoration degree information of the target object is adjusted; according to the first comparison result of the first frame rate and the target frame rate, or according to the preset frame rate threshold and the control point quantity corresponding relationship table, the control point quantity of the target object is adjusted; according to the first comparison result of the first frame rate and the target frame rate, or according to the preset frame rate threshold and the rendering resolution corresponding relationship table, the rendering resolution of the target object is adjusted.

[0213] Further, the processor is configured to:

[0214] Determine the reserved resource, the reserved resource is used to support the target object to output the second output result; adjust the display parameter of the target object according to the current idle resource and the reserved resource, so that the second frame rate of the second output result meets the target frame rate.

[0215] Further, the processor is further configured to:

[0216] Obtain the question information, the question information is the information to be answered obtained in the process of the target object outputting the target processing task processing result, and the question information includes audio information and image information; analyze the audio information and the image information respectively to obtain audio analysis results and image analysis results; determine the question text data based on the audio analysis results and the image analysis results; query the knowledge base based on the question text data, and generate answer information according to the query result of the knowledge base, so that the target object outputs the answer information.

[0217] Further, the processor is configured to:

[0218] Convert the audio information into first text data; perform lip analysis on a question object included in the image information to determine second text data corresponding to the lip shape of the question object, the question object being an object outputting the question information; perform semantic similarity analysis on the first text data and the second text data to obtain a similarity result; in a case where the similarity result represents that the semantic similarity between the first text data and the second text data reaches a similarity threshold, take any one of the first text data and the second text data as the question text data.

[0219] Further, the processor is configured to:

[0220] In a case where the similarity result characterizing the semantic similarity between the first text data and the second text data does not reach the similarity threshold, a first natural language fluency value of the first text data and a second natural language fluency value of the second text data are respectively determined; in a case where the first natural language fluency value is higher than the second natural language fluency value, the first text data is taken as the questioning text data; in a case where the first natural language fluency value is lower than the second natural language fluency value, the second text data is taken as the questioning text data.

[0221] The electronic device disclosed in the embodiment can be a first electronic device.

[0222] The electronic device disclosed in the embodiment is implemented based on the data processing method disclosed in the above embodiment, and thus will not be described herein.

[0223] The electronic device disclosed in the embodiment, after obtaining a target processing task including target multimedia data to be output, processes the target processing task to obtain a first output result, the target object is contained in the first output result, if the first frame rate of the first output result does not satisfy a target frame rate, the display parameter of the target object in the first output result needs to be adjusted to obtain a second output result, the second frame rate of the second output result satisfies the target frame rate, wherein the target object is a virtual image capable of outputting corresponding content of the target multimedia data. When the target processing task is determined and it is determined that the related content of the target multimedia data included in the target processing task needs to be output through the virtual image, the target processing task needs to be processed to obtain an output result, and when the frame rate of the output result does not satisfy the target frame rate, the display parameter of the virtual image in the output result is adjusted, so that the frame rate of the adjusted output result can satisfy the target frame rate, thereby avoiding the problem that the adjusted output result consumes high computing power when outputting the related content through the virtual image, and further affecting the output effect of the virtual image, and improving the user experience.

[0224] The embodiment of the present application further provides a readable storage medium, which has a computer program stored thereon, the computer program is loaded and executed by a processor, and each step of the above data processing method is implemented, and the specific implementation process can be referred to the description of the corresponding part of the above embodiment, and the embodiment will not be described herein.

[0225] The present application also proposes a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method provided in various optional implementation manners of the above data processing method, and the specific implementation process can be referred to the description of the corresponding embodiment, and will not be described herein.

[0226] It should be noted that the above-described apparatus embodiments are only illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection relationship between the modules in the apparatus embodiments provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0227] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, training device, or network device, etc.) execute the methods described in various embodiments of the present application.

[0228] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in the form of a computer program product in whole or in part.

[0229] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

Claims

1. A data processing method applied to a first electronic device, comprising: obtaining a target processing task, the target processing task comprising target multimedia data to be output; processing the target processing task to obtain a first output result comprising a target object; in a case where a first frame rate of the first output result does not satisfy a target frame rate, adjusting a display parameter of the target object, and outputting a second output result, a second frame rate of the second output result satisfying the target frame rate; wherein the target object is a virtual image capable of outputting target content, and the target content is content corresponding to the target multimedia data.

2. The method of claim 1, wherein the processing the target processing task to obtain the first output result comprising the target object comprises: analyzing the target processing task by a target model to obtain a knowledge base of the target processing task and commentary information of the target object in the target processing task; based on the knowledge base and the commentary information, obtaining the first output result comprising the target object.

3. The method of claim 2, wherein the analyzing the target processing task by the target model comprises at least one of: in a case where the target multimedia data comprises at least first image information, identifying content of the first image information by the target model to obtain the knowledge base and the commentary information; in a case where the target multimedia data comprises at least text information, converting the text information into second image information by the target model, and identifying content of the second image information to obtain the knowledge base and the commentary information; in a case where the target multimedia data comprises at least video information, identifying at least one video scene included in the video information by the target model, determining third image information corresponding to the at least one video scene in the video information, and identifying content of the third image information to obtain the knowledge base and the commentary information.

4. The method of claim 2, wherein the analyzing the target processing task by the target model further comprises: determining, by the target model, whether the target processing task comprises critical information; in a case where the target processing task comprises the critical information, processing the target processing task by the first electronic device; in a case where the target processing task does not comprise the critical information, processing the target processing task by the first electronic device and / or a second electronic device, a security level of the first electronic device being higher than that of the second electronic device, and the first electronic device and the second electronic device being data-connected.

5. The method of claim 4, wherein in the case where it is determined that the target processing task does not comprise the critical information, the method further comprises: in a case where it is determined that the first frame rate of the first output result output by the first electronic device does not satisfy the target frame rate, switching to processing the target processing task by the second electronic device; or ​ In a case where the first frame rate at which the second electronic device outputs the first output result does not satisfy the target frame rate, switching to processing of the target processing task by the first electronic device.

6. The method of any one of claims 1-5, wherein the adjusting the display parameter of the target object comprises at least one of the following: adjusting the modality information of the target object according to a preset frame rate threshold and modality information corresponding relationship table; adjusting the color restoration degree information of the target object according to the first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and color restoration degree information corresponding relationship table; adjusting the number of control points of the target object according to the first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and number of control points corresponding relationship table; adjusting the rendering resolution of the target object according to the first comparison result of the first frame rate and the target frame rate, or according to a preset frame rate threshold and rendering resolution corresponding relationship table.

7. The method of claim 2, further comprising: obtaining question information, the question information being information to be answered obtained in a process in which the target object outputs a processing result of the target processing task, the question information comprising audio information and image information; analyzing the audio information and the image information respectively to obtain audio analysis results and image analysis results; determining question text data based on the audio analysis results and the image analysis results; querying the knowledge base based on the question text data, and generating answer information according to a query result of the knowledge base, so that the target object outputs the answer information.

8. The method of claim 7, wherein the analyzing the audio information and the image information respectively to obtain audio analysis results and image analysis results comprises: converting the audio information into first text data; performing lip analysis on a question object included in the image information to determine second text data corresponding to a lip shape of the question object, the question object being an object outputting the question information; the determining question text data based on the audio analysis results and the image analysis results comprises: performing semantic similarity analysis on the first text data and the second text data to obtain a similarity result; in a case where the similarity result represents that a semantic similarity between the first text data and the second text data reaches a similarity threshold, taking any one of the first text data and the second text data as the question text data.

9. The method of claim 8, wherein the performing semantic similarity analysis on the first text data and the second text data to obtain a similarity result comprises: in a case where the similarity result represents that a semantic similarity between the first text data and the second text data does not reach a similarity threshold, respectively determining a first natural language fluency value of the first text data and a second natural language fluency value of the second text data. In a case where the first natural language fluency value is higher than the second natural language fluency value, the first text data is taken as the question text data; In a case where the first natural language fluency value is lower than the second natural language fluency value, the second text data is taken as the question text data.

10. An electronic device, comprising: a processor configured to obtain a target processing task, the target processing task comprising target multimedia data to be output; process the target processing task to obtain a first output result containing a target object; in a case where a first frame rate of the first output result does not meet a target frame rate, adjust a display parameter of the target object, and output a second output result, a second frame rate of the second output result meeting the target frame rate; wherein the target object is a virtual image capable of outputting target content, and the target content is content corresponding to the target multimedia data; a memory configured to store programs required by the processor to perform the above processing process.