Multimodal aigc generation service control method and system

By selecting a set of multimodal components from the AIGC library and loading and matching them based on user-end state information, the complexity of generating AI content of different modalities is solved, and an efficient and reliable multimodal AIGC service is achieved.

CN119759446BActive Publication Date: 2025-12-12HUIZHIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411896697.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-12-12
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing AIGC technology cannot achieve integrated and collaborative modes of different modalities such as text, images, video, and voice, resulting in complex and inefficient AI content generation.

Method used

Select a set of multimodal components from the AIGC library, load and match them based on the real-time working status information of the user, dynamically call up processing components, adjust the execution status, and integrate and save them when the generated task is correct.

Benefits of technology

It enables flexible and rapid generation of multimodal AIGC services, improves the reliability and efficiency of AI content generation, reduces generation complexity, and supports one-stop content generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759446B_ABST
    Figure CN119759446B_ABST
Patent Text Reader

Abstract

The application provides a multimodal AIGC generation service control method and system, selects processing components for different types of content from an AIGC library, and groups all the selected processing components into a multimodal component set, and loads the multimodal component set to a user end based on real-time working state information of the user end, so that the user end can perform multimodal AIGC services; based on the execution state information of the AI content generation task of the user end, the corresponding processing components are called from the multimodal component set to complete the corresponding AI content generation thread, so that different modal processing components can be flexibly and quickly used in the AI content generation process, and the generation reliability and efficiency of the AI content are improved; when the AI content generation task is correctly executed, the final generated content of the AI content generation task is output, and the multimodal component set currently loaded by the user end is integrated and saved, the complexity of AI content generation is reduced, and one-stop AIGC content generation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to a multi-modal AIGC generation service control method and system. BACKGROUND

[0002] Artificial Intelligence Generated Content (AIGC) is usually based on a deep learning algorithm platform to perform single AIGC service, which cannot integrate different modalities such as text, images, videos and voice into a collaborative mode of AIGC service, so that in the process of generating corresponding AI content, the deep learning algorithm platform needs to load single modal AIGC service respectively, which increases the complexity of AI content generation and cannot realize one-stop AIGC content integration generation. SUMMARY

[0003] The purpose of the present application is to provide a multi-modal AIGC generation service control method and system, which selects processing components for different types of content from an AIGC library, and forms a multi-modal component set by combining all selected processing components, and loads the multi-modal component set to the user terminal based on real-time working state information of the user terminal, so that the user terminal can perform multi-modal AIGC service; based on the execution state information of the AI content generation task of the user terminal, the corresponding processing component is called from the multi-modal component set to complete the corresponding AI content generation thread, so that different modal processing components can be used flexibly and quickly in the AI content generation process, improving the reliability and efficiency of AI content generation; when the AI content generation task is executed correctly, the final generated content of the AI content generation task is output, and the multi-modal component set currently loaded by the user terminal is integrated and saved, facilitating the user terminal to subsequently reuse the multi-modal component set for AIGC service, reducing the complexity of AI content generation, and realizing one-stop AIGC content generation.

[0004] The present application is achieved by the following technical solutions:

[0005] The multi-modal AIGC generation service control method comprises:

[0006] Based on the content generation request from the user terminal, processing components for different types of content are selected from an AIGC library, and all selected processing components are combined to form a multi-modal component set; based on real-time working state information of the user terminal, the multi-modal component set is loaded to the user terminal, and the user terminal is subjected to multi-modal component thread matching operation;

[0007] obtaining execution state information of the AI content generation task of the user terminal, based on the execution state information, calling corresponding processing components from the multi-modal component set to complete corresponding AI content generation threads; based on the execution state information of the AI content generation thread, updating the corresponding processing components in the multi-modal component set to adjust the execution state of the AI content generation thread;

[0008] detecting the final generated content of the AI content generation task, and determining whether the AI content generation task is executed correctly; when the AI content generation task is executed correctly, outputting the final generated content, and integrating and saving the multi-modal component set currently loaded by the user terminal.

[0009] Optionally, based on the content generation request from the user terminal, selecting processing components for different types of content from the AIGC library, and grouping all selected processing components into a multi-modal component set; based on the real-time working state information of the user terminal, loading the multi-modal component set to the user terminal, and performing multi-modal component thread matching operation on the user terminal, including:

[0010] analyzing and processing the content generation request from the user terminal to determine the modal attribute information of the AI content currently expected to be generated by the user terminal; based on the modal attribute information, selecting processing components for at least two of the text modal, image modal, video modal and voice modal from the AIGC library, and grouping all selected processing components into a multi-modal component set;

[0011] based on the real-time working state information of all threads of the user terminal, determining the thread with the maximum idle computing power in the user terminal; based on the position information of the thread with the maximum idle computing power in the user terminal, loading the multi-modal component set to the user terminal, and matching all processing components under the multi-modal component set to the corresponding link of the thread with the maximum idle computing power.

[0012] Optionally, obtaining execution state information of the AI content generation task of the user terminal, based on the execution state information, calling corresponding processing components from the multi-modal component set to complete corresponding AI content generation threads; based on the execution state information of the AI content generation thread, updating the corresponding processing components in the multi-modal component set to adjust the execution state of the AI content generation thread, including:

[0013] Obtain the execution progress information of the AI ​​content generation sub-task queue corresponding to the AI ​​content generation task on the user end, thereby determining the AI ​​content generation sub-task currently being executed on the user end; based on the currently executed AI content generation sub-task, retrieve the corresponding processing component from the multimodal component set to complete the AI ​​content generation thread corresponding to the AI ​​content generation sub-task;

[0014] The execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset execution speed threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread.

[0015] Optionally, based on the relationship between the execution speed and a preset execution speed threshold, anomaly determination is performed on the operation of the AI ​​content generation thread, including:

[0016] If the execution speed is not less than a preset execution speed threshold, the first running parameters of the AI ​​content generation thread are monitored in real time. The first running parameters include the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate.

[0017] The first operation evaluation coefficient is obtained by using the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate included in the first operation parameters.

[0018] The first operational evaluation coefficient is obtained using the following formula:

[0019]

[0020] Among them, R 01 denoted as the first performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi M represents the average difference in data volume between the input and output data corresponding to the i-th unit of time; i E represents the memory utilization rate corresponding to the i-th unit of time; i This represents the bit error rate of content generation corresponding to the i-th unit of time;

[0021] The first performance evaluation coefficient is compared with a preset coefficient threshold.

[0022] When the first performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation and an abnormal risk warning is issued.

[0023] If the execution speed is less than a preset execution speed threshold, the second running parameters of the AI ​​content generation thread are monitored in real time. The second running parameters include the amount of content data generated per unit time, the difference in data volume between the input content and the output content, and the execution speed value.

[0024] The second operation evaluation coefficient is obtained by using the amount of content data generated per unit time, the difference in data volume between input content and output content, and the execution speed value contained in the second operation parameters.

[0025] The second operational evaluation coefficient is obtained using the following formula:

[0026]

[0027] Among them, R 02 denoted as the second performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi C represents the average difference between the input and output data at the i-th unit of time; zmaxi This represents the maximum difference in data volume between the input and output data corresponding to the i-th unit of time; v i This represents the execution speed value generated in the i-th unit of time; v b This represents the standard deviation of the generated execution speed value corresponding to n units of time;

[0028] The second performance evaluation coefficient is compared with a preset coefficient threshold.

[0029] When the second performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation, and an abnormal risk warning is issued.

[0030] Optionally, the final generated content of the AI ​​content generation task is detected to determine whether the AI ​​content generation task has been executed correctly; if the AI ​​content generation task has been executed correctly, the final generated content is output, and the multimodal component set currently loaded on the user terminal is integrated and saved, including:

[0031] The final generated content of the AI ​​content generation task is detected to obtain the data error rate of the final generated content; if the data error rate is greater than or equal to a preset error rate threshold, it is determined that the AI ​​content generation task has not been executed correctly; otherwise, it is determined that the AI ​​content generation task has been executed correctly.

[0032] When the AI ​​content generation task is executed correctly, the final generated content is backed up and then output. The set of multimodal components currently loaded on the user terminal is integrated into a set of reusable multimodal components and stored in a designated storage space within the user terminal.

[0033] The multimodal AIGC generation service control system includes:

[0034] The multimodal component determination module is used to select processing components for different types of content from the AIGC library based on the content generation request from the user, and to form a multimodal component set by combining all the selected processing components.

[0035] The multimodal component matching module is used to load the multimodal component set into the user terminal based on the real-time working status information of the user terminal, and to perform multimodal component thread matching operation on the user terminal.

[0036] The AI ​​content generation task processing module is used to obtain the execution status information of the AI ​​content generation task on the user end, and based on the execution status information, to retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; and based on the execution status information of the AI ​​content generation thread, to update the corresponding processing component in the multimodal component set, thereby adjusting the execution status of the AI ​​content generation thread.

[0037] The AI ​​content generation task judgment module is used to detect the final generated content of the AI ​​content generation task and determine whether the AI ​​content generation task is executed correctly.

[0038] The multimodal component set operation module is used to output the final generated content when the AI ​​content generation task is executed correctly, and to integrate and save the multimodal component set currently loaded on the user terminal.

[0039] Optionally, the multimodal component determination module is used to select processing components for different types of content from the AIGC library based on content generation requests from the user terminal, and to assemble all selected processing components into a multimodal component set, including:

[0040] The content generation request from the user terminal is parsed and processed to determine the modal attribute information of the AI ​​content that the user terminal currently expects to generate; based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set;

[0041] The multimodal component matching module is used to load the multimodal component set onto the user terminal based on the real-time working status information of the user terminal, and to perform multimodal component thread matching operations on the user terminal, including:

[0042] Based on the real-time working status information of all threads under the user terminal, the thread with the maximum idle computing power in the user terminal is determined; based on the location information of the thread with the maximum idle computing power in the user terminal, the multimodal component set is loaded into the user terminal, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power.

[0043] Optionally, the AI ​​content generation task processing module is used to obtain the execution status information of the AI ​​content generation task on the user end, and based on the execution status information, to retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; based on the execution status information of the AI ​​content generation thread, to update the corresponding processing component in the multimodal component set, thereby adjusting the execution status of the AI ​​content generation thread, including:

[0044] Obtain the execution progress information of the AI ​​content generation sub-task queue corresponding to the AI ​​content generation task on the user end, thereby determining the AI ​​content generation sub-task currently being executed on the user end; based on the currently executed AI content generation sub-task, retrieve the corresponding processing component from the multimodal component set to complete the AI ​​content generation thread corresponding to the AI ​​content generation sub-task;

[0045] The execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset execution speed threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread.

[0046] Optionally, based on the relationship between the execution speed and a preset execution speed threshold, anomaly determination is performed on the operation of the AI ​​content generation thread, including:

[0047] If the execution speed is not less than a preset execution speed threshold, the first running parameters of the AI ​​content generation thread are monitored in real time. The first running parameters include the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate.

[0048] The first operation evaluation coefficient is obtained by using the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate included in the first operation parameters.

[0049] The first operational evaluation coefficient is obtained using the following formula:

[0050]

[0051] Among them, R 01 denoted as the first performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi M represents the average difference in data volume between the input and output data corresponding to the i-th unit of time; i E represents the memory utilization rate corresponding to the i-th unit of time; i This represents the bit error rate of content generation corresponding to the i-th unit of time;

[0052] The first performance evaluation coefficient is compared with a preset coefficient threshold.

[0053] When the first performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation and an abnormal risk warning is issued.

[0054] If the execution speed is less than a preset execution speed threshold, the second running parameters of the AI ​​content generation thread are monitored in real time. The second running parameters include the amount of content data generated per unit time, the difference in data volume between the input content and the output content, and the execution speed value.

[0055] The second operation evaluation coefficient is obtained by using the amount of content data generated per unit time, the difference in data volume between input content and output content, and the execution speed value contained in the second operation parameters.

[0056] The second operational evaluation coefficient is obtained using the following formula:

[0057]

[0058] Among them, R 02 denoted as the second performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi C represents the average difference between the input and output data at the i-th unit of time; zmaxi This represents the maximum difference in data volume between the input and output data corresponding to the i-th unit of time; v i This represents the execution speed value generated in the i-th unit of time; v b This represents the standard deviation of the generated execution speed value corresponding to n units of time;

[0059] The second performance evaluation coefficient is compared with a preset coefficient threshold.

[0060] When the second performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation, and an abnormal risk warning is issued.

[0061] Optionally, the AI ​​content generation task judgment module is used to detect the final generated content of the AI ​​content generation task and determine whether the AI ​​content generation task has been executed correctly, including:

[0062] The final generated content of the AI ​​content generation task is detected to obtain the data error rate of the final generated content; if the data error rate is greater than or equal to a preset error rate threshold, it is determined that the AI ​​content generation task has not been executed correctly; otherwise, it is determined that the AI ​​content generation task has been executed correctly.

[0063] The multimodal component set operation module is used to output the final generated content when the AI ​​content generation task is executed correctly, and to integrate and save the multimodal component set currently loaded on the user terminal, including:

[0064] When the AI ​​content generation task is executed correctly, the final generated content is backed up and then output. The set of multimodal components currently loaded on the user terminal is integrated into a set of reusable multimodal components and stored in a designated storage space within the user terminal.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] The multimodal AIGC generation service control method and system provided in this application selects processing components for different types of content from the AIGC library, and assembles all selected processing components into a multimodal component set. Based on the real-time working status information of the user terminal, the multimodal component set is loaded onto the user terminal, enabling the user terminal to perform multimodal AIGC services. Based on the execution status information of the AI ​​content generation task on the user terminal, the corresponding processing components are retrieved from the multimodal component set to complete the corresponding AI content generation thread. This allows for flexible and rapid use of different modal processing components during the AI ​​content generation process, improving the reliability and efficiency of AI content generation. When the AI ​​content generation task is executed correctly, the final generated content of the AI ​​content generation task is output, and the multimodal component set currently loaded on the user terminal is integrated and saved. This facilitates the user terminal to reuse the multimodal component set for AIGC services in the future, reducing the complexity of AI content generation and achieving one-stop AIGC content generation. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0068] Figure 1 This is a flowchart illustrating the multimodal AIGC generation service control method provided by the present invention.

[0069] Figure 2 This is a schematic diagram of the structure of the multimodal AIGC generation service control system provided by the present invention. Detailed Implementation

[0070] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, it should be noted that, for ease of description, only the parts relevant to this application are shown in the accompanying drawings, not the entire structure. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0071] The terms “comprising” and “having”, and any variations thereof, used in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0072] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0073] Please see Figure 1 As shown, an embodiment of this application provides a multimodal AIGC generation service control method including:

[0074] Based on the content generation request from the user, processing components for different types of content are selected from the AIGC library, and all selected processing components are combined into a multimodal component set; based on the real-time working status information of the user, the multimodal component set is loaded into the user and a multimodal component thread matching operation is performed on the user.

[0075] The system obtains the execution status information of the AI ​​content generation task on the user end. Based on the execution status information, it retrieves the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread. Based on the execution status information of the AI ​​content generation thread, it updates the corresponding processing component in the multimodal component set to adjust the execution status of the AI ​​content generation thread.

[0076] The final generated content of the AI ​​content generation task is detected to determine whether the AI ​​content generation task was executed correctly. If the AI ​​content generation task was executed correctly, the final generated content is output, and the collection of multimodal components currently loaded on the user terminal is integrated and saved.

[0077] The beneficial effects of the above embodiments are as follows: the multimodal AIGC generation service control method selects processing components for different types of content from the AIGC library, assembles all selected processing components into a multimodal component set, and loads the multimodal component set onto the user terminal based on the real-time working status information of the user terminal, enabling the user terminal to perform multimodal AIGC services; based on the execution status information of the AI ​​content generation task on the user terminal, the corresponding processing components are retrieved from the multimodal component set to complete the corresponding AI content generation thread, enabling flexible and rapid use of different modal processing components during the AI ​​content generation process, improving the reliability and efficiency of AI content generation; when the AI ​​content generation task is executed correctly, the final generated content of the AI ​​content generation task is output, and the multimodal component set currently loaded on the user terminal is integrated and saved, facilitating the user terminal to reuse the multimodal component set for AIGC services in the future, reducing the complexity of AI content generation, and realizing one-stop AIGC content generation.

[0078] In another embodiment, based on a content generation request from the user client, processing components for different types of content are selected from the AIGC library, and all selected processing components are combined into a multimodal component set; based on the real-time working status information of the user client, the multimodal component set is loaded onto the user client, and a multimodal component thread matching operation is performed on the user client, including:

[0079] The content generation request from the user terminal is parsed and processed to determine the modal attribute information of the AI ​​content that the user terminal currently expects to generate; based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set;

[0080] Based on the real-time working status information of all threads under the client, the thread with the maximum idle computing power in the client is determined; based on the location information of the thread with the maximum idle computing power in the client, the multimodal component set is loaded into the client, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power.

[0081] The beneficial effect of the above embodiments is that when a user needs to obtain AIGC services to generate corresponding AI content, the content generation request from the user is first parsed to determine the modal attribute information of the AI ​​content that the user currently expects to generate. This modal attribute information may include, but is not limited to, text modal attribute information, image modal attribute information, video modal attribute information, and speech modal attribute information, thus enabling comprehensive modal recognition of the AI ​​content that the user currently expects to generate. Then, based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality, and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set, so that the formed multimodal component set can provide the user with a matching multimodal AIGC service. To ensure the proper formation of the multimodal component set within the user client, it is necessary to run the multimodal component set using threads within the user client. At this time, based on the real-time working status information of all threads under the user client, the thread with the maximum idle computing power in the user client is determined. Then, according to the location information of the thread with the maximum idle computing power in the user client, the multimodal component set is loaded into the user client, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power, so that the multimodal component set can be fully triggered and run within the user client.

[0082] In another embodiment, the execution status information of the AI ​​content generation task on the user terminal is obtained. Based on the execution status information, the corresponding processing component is retrieved from the multimodal component set to complete the corresponding AI content generation thread. Based on the execution status information of the AI ​​content generation thread, the corresponding processing component in the multimodal component set is updated to adjust the execution status of the AI ​​content generation thread, including:

[0083] Obtain the execution progress information of the AI ​​content generation subtask queue corresponding to the AI ​​content generation task of the user terminal, so as to determine the AI ​​content generation subtask currently being executed by the user terminal; based on the currently executed AI content generation subtask, call the corresponding processing component from the multimodal component set to complete the AI ​​content generation thread corresponding to the AI ​​content generation subtask;

[0084] The execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset execution speed threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread.

[0085] The beneficial effects of the above embodiments are that by obtaining the execution progress information of the AI ​​content generation sub-task queue corresponding to the AI ​​content generation task on the user terminal, the current AI content generation sub-task being executed on the user terminal can be determined. The corresponding processing component is then retrieved from the multimodal component set to complete the AI ​​content generation thread corresponding to the sub-task, thus enabling accurate determination of the AI ​​content generation status within the user terminal. Furthermore, the execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread and improve the reliability of AI content generation on the user terminal.

[0086] In another embodiment, based on the relationship between the execution speed and a preset execution speed threshold, anomaly determination is performed on the operation of the AI ​​content generation thread, including:

[0087] If the execution speed is not less than a preset execution speed threshold, the first running parameters of the AI ​​content generation thread are monitored in real time. The first running parameters include the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate.

[0088] The first operation evaluation coefficient is obtained by using the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate included in the first operation parameters.

[0089] The first operational evaluation coefficient is obtained using the following formula:

[0090]

[0091] Among them, R 01 denoted as the first performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C iC represents the amount of content data generated in the i-th unit of time; zpi M represents the average difference in data volume between the input and output data corresponding to the i-th unit of time; i E represents the memory utilization rate corresponding to the i-th unit of time; i This represents the bit error rate of content generation corresponding to the i-th unit of time;

[0092] The first performance evaluation coefficient is compared with a preset coefficient threshold.

[0093] When the first performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation and an abnormal risk warning is issued.

[0094] If the execution speed is less than a preset execution speed threshold, the second running parameters of the AI ​​content generation thread are monitored in real time. The second running parameters include the amount of content data generated per unit time, the difference in data volume between the input content and the output content, and the execution speed value.

[0095] The second operation evaluation coefficient is obtained by using the amount of content data generated per unit time, the difference in data volume between input content and output content, and the execution speed value contained in the second operation parameters.

[0096] The second operational evaluation coefficient is obtained using the following formula:

[0097]

[0098] Among them, R 02 denoted as the second performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi C represents the average difference between the input and output data at the i-th unit of time; zmaxi This represents the maximum difference in data volume between the input and output data corresponding to the i-th unit of time; v i This represents the execution speed value generated in the i-th unit of time; v b This represents the standard deviation of the generated execution speed value corresponding to n units of time;

[0099] The second performance evaluation coefficient is compared with a preset coefficient threshold.

[0100] When the second performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation, and an abnormal risk warning is issued.

[0101] The beneficial effects of the above embodiments are that this scheme can dynamically select and monitor different operating parameters based on the relationship between the execution speed (i.e., processing efficiency) of the AI ​​content generation thread and a preset execution speed threshold. When the execution speed is not lower than the preset threshold, more comprehensive performance indicators (such as content data volume, memory utilization, data difference, and bit error rate) are considered; while when the execution speed is lower than the threshold, the focus is on generation efficiency and its related parameters (such as content data volume, data difference, and execution speed). This dynamic adjustment improves the targeting and efficiency of monitoring. By introducing a first operating evaluation coefficient (R01) and a second operating evaluation coefficient (R02), this scheme can comprehensively evaluate the operating status of the AI ​​content generation thread by integrating multiple key operating parameters. This comprehensive evaluation method is more accurate than single-parameter monitoring and can more comprehensively reflect the operating health of the thread. The coefficient calculation considers the changes in the time series (n units of time), which helps to capture the trend of performance fluctuations and thus detect potential operating anomalies in advance. When the operating evaluation coefficient exceeds the preset coefficient threshold, the system will immediately determine that the AI ​​content generation thread has a risk of operating anomalies and trigger an early warning mechanism. This real-time anomaly detection and early warning capability facilitates rapid response and problem-solving, preventing adverse consequences such as performance degradation or system crashes. Through continuous monitoring and anomaly identification, this solution helps to promptly identify and resolve performance bottlenecks or improper resource allocation issues in the AI ​​content generation thread. This not only improves the overall system efficiency but also optimizes resource utilization, enhancing system stability and reliability. The aforementioned technical solution has broad applicability and can be easily extended to other types of thread or process monitoring, providing strong support for system performance management and optimization.

[0102] In summary, this technical solution achieves precise monitoring and anomaly detection of the AI ​​content generation thread's running status through dynamic monitoring, comprehensive evaluation, and real-time early warning, providing an effective technical means to improve system performance, optimize resource management, and ensure system stability.

[0103] In another embodiment, the final generated content of the AI ​​content generation task is detected to determine whether the AI ​​content generation task has been executed correctly; if the AI ​​content generation task has been executed correctly, the final generated content is output, and the integration and saving operation of the multimodal component set currently loaded on the user terminal is performed, including:

[0104] The final generated content of the AI ​​content generation task is inspected to obtain the bit error rate of the final generated content; if the bit error rate is greater than or equal to the preset bit error rate threshold, it is determined that the AI ​​content generation task has not been executed correctly; otherwise, it is determined that the AI ​​content generation task has been executed correctly.

[0105] When the AI ​​content generation task is executed correctly, the final generated content is backed up and then output. The set of multimodal components currently loaded on the client is integrated into a set of reusable multimodal components and stored in the specified storage space inside the client.

[0106] The beneficial effects of the above embodiments are as follows: The final generated content of the AI ​​content generation task is detected to obtain the data error rate of the final generated content, thereby accurately determining whether the AI ​​content generation task on the user end is executed correctly; furthermore, when the AI ​​content generation task is executed correctly, the final generated content is backed up before being output; and the currently loaded multimodal component set on the user end is integrated into a reusable multimodal component set and stored in a designated storage space within the user end, ensuring the normal display of the final generated content. This facilitates the user end's subsequent reusability of the multimodal component set for AIGC services, reduces the complexity of AI content generation, and achieves one-stop AIGC content generation.

[0107] Please see Figure 2 As shown, an embodiment of this application provides a multimodal AIGC generation service control system including:

[0108] The multimodal component determination module is used to select processing components for different types of content from the AIGC library based on the content generation request from the user, and to form a multimodal component set by combining all the selected processing components.

[0109] The multimodal component matching module is used to load the multimodal component set into the user terminal based on the real-time working status information of the user terminal, and to perform multimodal component thread matching operation on the user terminal.

[0110] The AI ​​content generation task processing module is used to obtain the execution status information of the AI ​​content generation task on the user end, and based on the execution status information, to retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; based on the execution status information of the AI ​​content generation thread, to update the corresponding processing component in the multimodal component set, thereby adjusting the execution status of the AI ​​content generation thread.

[0111] The AI ​​content generation task judgment module is used to detect the final generated content of the AI ​​content generation task and determine whether the AI ​​content generation task has been executed correctly.

[0112] The multimodal component collection operation module is used to output the final generated content when the AI ​​content generation task is executed correctly, and to integrate and save the multimodal component collection currently loaded on the user's end.

[0113] The beneficial effects of the above embodiments are as follows: the multimodal AIGC generation service control system selects processing components for different types of content from the AIGC library, assembles all selected processing components into a multimodal component set, and loads the multimodal component set onto the user terminal based on the real-time working status information of the user terminal, enabling the user terminal to perform multimodal AIGC services; based on the execution status information of the AI ​​content generation task on the user terminal, it retrieves the corresponding processing components from the multimodal component set to complete the corresponding AI content generation thread, enabling flexible and rapid use of different modal processing components during the AI ​​content generation process, improving the reliability and efficiency of AI content generation; when the AI ​​content generation task is executed correctly, it outputs the final generated content of the AI ​​content generation task, and integrates and saves the multimodal component set currently loaded on the user terminal, facilitating the user terminal to reuse the multimodal component set for AIGC services in the future, reducing the complexity of AI content generation, and realizing one-stop AIGC content generation.

[0114] In another embodiment, the multimodal component determination module is used to select processing components for different types of content from the AIGC library based on content generation requests from the user, and to assemble all selected processing components into a multimodal component set, including:

[0115] The content generation request from the user terminal is parsed and processed to determine the modal attribute information of the AI ​​content that the user terminal currently expects to generate; based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set;

[0116] The multimodal component matching module is used to load the multimodal component set onto the user terminal based on the real-time working status information of the user terminal, and to perform multimodal component thread matching operations on the user terminal, including:

[0117] Based on the real-time working status information of all threads under the client, the thread with the maximum idle computing power in the client is determined; based on the location information of the thread with the maximum idle computing power in the client, the multimodal component set is loaded into the client, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power.

[0118] The beneficial effect of the above embodiments is that when a user needs to obtain AIGC services to generate corresponding AI content, the content generation request from the user is first parsed to determine the modal attribute information of the AI ​​content that the user currently expects to generate. This modal attribute information may include, but is not limited to, text modal attribute information, image modal attribute information, video modal attribute information, and speech modal attribute information, thus enabling comprehensive modal recognition of the AI ​​content that the user currently expects to generate. Then, based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality, and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set, so that the formed multimodal component set can provide the user with a matching multimodal AIGC service. To ensure the proper formation of the multimodal component set within the user client, it is necessary to run the multimodal component set using threads within the user client. At this time, based on the real-time working status information of all threads under the user client, the thread with the maximum idle computing power in the user client is determined. Then, according to the location information of the thread with the maximum idle computing power in the user client, the multimodal component set is loaded into the user client, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power, so that the multimodal component set can be fully triggered and run within the user client.

[0119] In another embodiment, the AI ​​content generation task processing module is used to obtain the execution status information of the AI ​​content generation task on the user end, and based on the execution status information, to retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; based on the execution status information of the AI ​​content generation thread, to update the corresponding processing component in the multimodal component set, thereby adjusting the execution status of the AI ​​content generation thread, including:

[0120] Obtain the execution progress information of the AI ​​content generation subtask queue corresponding to the AI ​​content generation task of the user terminal, so as to determine the AI ​​content generation subtask currently being executed by the user terminal; based on the currently executed AI content generation subtask, call the corresponding processing component from the multimodal component set to complete the AI ​​content generation thread corresponding to the AI ​​content generation subtask;

[0121] The execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset execution speed threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread.

[0122] The beneficial effects of the above embodiments are that by obtaining the execution progress information of the AI ​​content generation sub-task queue corresponding to the AI ​​content generation task on the user terminal, the current AI content generation sub-task being executed on the user terminal can be determined. The corresponding processing component is then retrieved from the multimodal component set to complete the AI ​​content generation thread corresponding to the sub-task, thus enabling accurate determination of the AI ​​content generation status within the user terminal. Furthermore, the execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread and improve the reliability of AI content generation on the user terminal.

[0123] In another embodiment, based on the relationship between the execution speed and a preset execution speed threshold, anomaly determination is performed on the operation of the AI ​​content generation thread, including:

[0124] If the execution speed is not less than a preset execution speed threshold, the first running parameters of the AI ​​content generation thread are monitored in real time. The first running parameters include the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate.

[0125] The first operation evaluation coefficient is obtained by using the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate included in the first operation parameters.

[0126] The first operational evaluation coefficient is obtained using the following formula:

[0127]

[0128] Among them, R 01 denoted as the first performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi M represents the average difference in data volume between the input and output data corresponding to the i-th unit of time; i E represents the memory utilization rate corresponding to the i-th unit of time; i This represents the bit error rate of content generation corresponding to the i-th unit of time;

[0129] The first performance evaluation coefficient is compared with a preset coefficient threshold.

[0130] When the first performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation and an abnormal risk warning is issued.

[0131] If the execution speed is less than a preset execution speed threshold, the second running parameters of the AI ​​content generation thread are monitored in real time. The second running parameters include the amount of content data generated per unit time, the difference in data volume between the input content and the output content, and the execution speed value.

[0132] The second operation evaluation coefficient is obtained by using the amount of content data generated per unit time, the difference in data volume between input content and output content, and the execution speed value contained in the second operation parameters.

[0133] The second operational evaluation coefficient is obtained using the following formula:

[0134]

[0135] Among them, R 02 denoted as the second performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi C represents the average difference between the input and output data at the i-th unit of time; zmaxi This represents the maximum difference in data volume between the input and output data corresponding to the i-th unit of time; v i This represents the execution speed value generated in the i-th unit of time; v b This represents the standard deviation of the generated execution speed value corresponding to n units of time;

[0136] The second performance evaluation coefficient is compared with a preset coefficient threshold.

[0137] When the second performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation, and an abnormal risk warning is issued.

[0138] The beneficial effects of the above embodiments are that this scheme can dynamically select and monitor different operating parameters based on the relationship between the execution speed (i.e., processing efficiency) of the AI ​​content generation thread and a preset execution speed threshold. When the execution speed is not lower than the preset threshold, more comprehensive performance indicators (such as content data volume, memory utilization, data difference, and bit error rate) are considered; while when the execution speed is lower than the threshold, the focus is on generation efficiency and its related parameters (such as content data volume, data difference, and execution speed). This dynamic adjustment improves the targeting and efficiency of monitoring. By introducing a first operating evaluation coefficient (R01) and a second operating evaluation coefficient (R02), this scheme can comprehensively evaluate the operating status of the AI ​​content generation thread by integrating multiple key operating parameters. This comprehensive evaluation method is more accurate than single-parameter monitoring and can more comprehensively reflect the operating health of the thread. The coefficient calculation considers the changes in the time series (n units of time), which helps to capture the trend of performance fluctuations and thus detect potential operating anomalies in advance. When the operating evaluation coefficient exceeds the preset coefficient threshold, the system will immediately determine that the AI ​​content generation thread has a risk of operating anomalies and trigger an early warning mechanism. This real-time anomaly detection and early warning capability facilitates rapid response and problem-solving, preventing adverse consequences such as performance degradation or system crashes. Through continuous monitoring and anomaly identification, this solution helps to promptly identify and resolve performance bottlenecks or improper resource allocation issues in the AI ​​content generation thread. This not only improves the overall system efficiency but also optimizes resource utilization, enhancing system stability and reliability. The aforementioned technical solution has broad applicability and can be easily extended to other types of thread or process monitoring, providing strong support for system performance management and optimization.

[0139] In summary, this technical solution achieves precise monitoring and anomaly detection of the AI ​​content generation thread's running status through dynamic monitoring, comprehensive evaluation, and real-time early warning, providing an effective technical means to improve system performance, optimize resource management, and ensure system stability.

[0140] In another embodiment, the AI ​​content generation task judgment module is used to detect the final generated content of the AI ​​content generation task and determine whether the AI ​​content generation task has been executed correctly, including:

[0141] The final generated content of the AI ​​content generation task is inspected to obtain the bit error rate of the final generated content; if the bit error rate is greater than or equal to the preset bit error rate threshold, it is determined that the AI ​​content generation task has not been executed correctly; otherwise, it is determined that the AI ​​content generation task has been executed correctly.

[0142] This multimodal component set operation module is used to output the final generated content when the AI ​​content generation task is executed correctly, and to integrate and save the multimodal component set currently loaded on the user's end, including:

[0143] When the AI ​​content generation task is executed correctly, the final generated content is backed up and then output. The set of multimodal components currently loaded on the client is integrated into a set of reusable multimodal components and stored in the specified storage space inside the client.

[0144] The beneficial effects of the above embodiments are as follows: The final generated content of the AI ​​content generation task is detected to obtain the data error rate of the final generated content, thereby accurately determining whether the AI ​​content generation task on the user end is executed correctly; furthermore, when the AI ​​content generation task is executed correctly, the final generated content is backed up before being output; and the currently loaded multimodal component set on the user end is integrated into a reusable multimodal component set and stored in a designated storage space within the user end, ensuring the normal display of the final generated content. This facilitates the user end's subsequent reusability of the multimodal component set for AIGC services, reduces the complexity of AI content generation, and achieves one-stop AIGC content generation.

[0145] In summary, this multimodal AIGC generation service control method and system selects processing components for different types of content from the AIGC library, assembles all selected processing components into a multimodal component set, and loads the multimodal component set onto the user terminal based on the user's real-time working status information, enabling the user terminal to perform multimodal AIGC services. Based on the execution status information of the user terminal's AI content generation task, the system retrieves the corresponding processing components from the multimodal component set to complete the corresponding AI content generation thread, allowing for flexible and rapid use of different modal processing components during the AI ​​content generation process, improving the reliability and efficiency of AI content generation. When the AI ​​content generation task is executed correctly, the system outputs the final generated content of the AI ​​content generation task and integrates and saves the multimodal component set currently loaded on the user terminal, facilitating the user terminal's ability to reuse the multimodal component set for AIGC services in the future, reducing the complexity of AI content generation, and achieving one-stop AIGC content generation.

[0146] The above is only one specific embodiment of the present invention, and any improvements made based on the concept of the present invention shall be considered within the scope of protection of the present invention.

Claims

1. A multimodal AIGC generation service control method, characterized in that, include: Based on the content generation request from the user, processing components for different types of content are selected from the AIGC library, and all selected processing components are combined into a multimodal component set; based on the real-time working status information of the user, the multimodal component set is loaded into the user and a multimodal component thread matching operation is performed on the user; The execution status information of the AI ​​content generation task on the user terminal is obtained. Based on the execution status information, the corresponding processing component is retrieved from the multimodal component set to complete the corresponding AI content generation thread. Based on the execution status information of the AI ​​content generation thread, the corresponding processing component in the multimodal component set is updated to adjust the execution status of the AI ​​content generation thread. The final generated content of the AI ​​content generation task is detected to determine whether the AI ​​content generation task is executed correctly; if the AI ​​content generation task is executed correctly, the final generated content is output, and the multimodal component set currently loaded on the user terminal is integrated and saved. Specifically, based on the real-time working status information of the user terminal, the multimodal component set is loaded onto the user terminal, and a multimodal component thread matching operation is performed on the user terminal, including: Based on the real-time working status information of all threads under the user terminal, the thread with the maximum idle computing power in the user terminal is determined; based on the location information of the thread with the maximum idle computing power in the user terminal, the multimodal component set is loaded into the user terminal, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power.

2. The multimodal AIGC generation service control method as described in claim 1, characterized in that: Based on content generation requests from the user, processing components for different types of content are selected from the AIGC library, and all selected processing components are combined into a multimodal component set. The content generation request from the user terminal is parsed and processed to determine the modal attribute information of the AI ​​content that the user terminal currently expects to generate; based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set.

3. The multimodal AIGC generation service control method as described in claim 1, characterized in that: Obtain the execution status information of the AI ​​content generation task on the user end; based on the execution status information, retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; based on the execution status information of the AI ​​content generation thread, update the corresponding processing component in the multimodal component set to adjust the execution status of the AI ​​content generation thread, including: Obtain the execution progress information of the AI ​​content generation sub-task queue corresponding to the AI ​​content generation task on the user end, thereby determining the AI ​​content generation sub-task currently being executed on the user end; based on the currently executed AI content generation sub-task, retrieve the corresponding processing component from the multimodal component set to complete the AI ​​content generation thread corresponding to the AI ​​content generation sub-task; The execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset execution speed threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread.

4. The multimodal AIGC generation service control method as described in claim 3, characterized in that, The multimodal AIGC generation service control method also includes: If the execution speed is not less than a preset execution speed threshold, the first running parameters of the AI ​​content generation thread are monitored in real time. The first running parameters include the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate. The first operation evaluation coefficient is obtained by using the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate included in the first operation parameters. The first operational evaluation coefficient is obtained using the following formula: Among them, R 01 denoted as the first performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi M represents the average difference in data volume between the input and output data corresponding to the i-th unit of time; i E represents the memory utilization rate corresponding to the i-th unit of time; i This represents the bit error rate of content generation corresponding to the i-th unit of time; The first performance evaluation coefficient is compared with a preset coefficient threshold. When the first performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation and an abnormal risk warning is issued. If the execution speed is less than a preset execution speed threshold, the second running parameters of the AI ​​content generation thread are monitored in real time. The second running parameters include the amount of content data generated per unit time, the difference in data volume between the input content and the output content, and the execution speed value. The second operation evaluation coefficient is obtained by using the amount of content data generated per unit time, the difference in data volume between input content and output content, and the execution speed value contained in the second operation parameters. The second operational evaluation coefficient is obtained using the following formula: Among them, R 02 denoted as the second performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi C represents the average difference between the input and output data at the i-th unit of time; zmaxi This represents the maximum difference in data volume between the input and output data corresponding to the i-th unit of time; v i This represents the execution speed value generated in the i-th unit of time; v b This represents the standard deviation of the generated execution speed value corresponding to n units of time; The second performance evaluation coefficient is compared with a preset coefficient threshold. When the second performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation, and an abnormal risk warning is issued.

5. The multimodal AIGC generation service control method as described in claim 1, characterized in that: The final generated content of the AI ​​content generation task is detected to determine whether the AI ​​content generation task was executed correctly; if the AI ​​content generation task was executed correctly, the final generated content is output, and the multimodal component set currently loaded on the user terminal is integrated and saved, including: The final generated content of the AI ​​content generation task is detected to obtain the data error rate of the final generated content; if the data error rate is greater than or equal to a preset error rate threshold, it is determined that the AI ​​content generation task has not been executed correctly; otherwise, it is determined that the AI ​​content generation task has been executed correctly. When the AI ​​content generation task is executed correctly, the final generated content is backed up and then output. The set of multimodal components currently loaded on the user terminal is integrated into a set of reusable multimodal components and stored in a designated storage space within the user terminal.

6. A multimodal AIGC generation service control system, characterized in that, include: The multimodal component determination module is used to select processing components for different types of content from the AIGC library based on the content generation request from the user, and to form a multimodal component set by combining all the selected processing components. The multimodal component matching module is used to load the multimodal component set into the user terminal based on the real-time working status information of the user terminal, and to perform multimodal component thread matching operation on the user terminal. The AI ​​content generation task processing module is used to obtain the execution status information of the AI ​​content generation task on the user end, and based on the execution status information, to retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; and based on the execution status information of the AI ​​content generation thread, to update the corresponding processing component in the multimodal component set, thereby adjusting the execution status of the AI ​​content generation thread. The AI ​​content generation task judgment module is used to detect the final generated content of the AI ​​content generation task and determine whether the AI ​​content generation task is executed correctly. The multimodal component set operation module is used to output the final generated content when the AI ​​content generation task is executed correctly, and to integrate and save the multimodal component set currently loaded on the user terminal. The multimodal component matching module is used to load the multimodal component set onto the user terminal based on the real-time working status information of the user terminal, and to perform multimodal component thread matching operations on the user terminal, including: Based on the real-time working status information of all threads under the user terminal, the thread with the maximum idle computing power in the user terminal is determined; based on the location information of the thread with the maximum idle computing power in the user terminal, the multimodal component set is loaded into the user terminal, and all processing components under the multimodal component set are matched to the corresponding links of the thread with the maximum idle computing power.

7. The multimodal AIGC generation service control system as described in claim 6, characterized in that: The multimodal component determination module is used to select processing components for different types of content from the AIGC library based on content generation requests from the user terminal, and to assemble all selected processing components into a multimodal component set, including: The content generation request from the user terminal is parsed and processed to determine the modal attribute information of the AI ​​content that the user terminal currently expects to generate; based on the modal attribute information, processing components for at least two of the text modality, image modality, video modality and speech modality are selected from the AIGC library, and all selected processing components are combined into a multimodal component set.

8. The multimodal AIGC generation service control system as described in claim 6, characterized in that: The AI ​​content generation task processing module is used to obtain the execution status information of the AI ​​content generation task on the user end, and based on the execution status information, to retrieve the corresponding processing component from the multimodal component set to complete the corresponding AI content generation thread; based on the execution status information of the AI ​​content generation thread, to update the corresponding processing component in the multimodal component set, thereby adjusting the execution status of the AI ​​content generation thread, including: Obtain the execution progress information of the AI ​​content generation sub-task queue corresponding to the AI ​​content generation task on the user end, thereby determining the AI ​​content generation sub-task currently being executed on the user end; based on the currently executed AI content generation sub-task, retrieve the corresponding processing component from the multimodal component set to complete the AI ​​content generation thread corresponding to the AI ​​content generation sub-task; The execution speed of the AI ​​content generation thread is compared with a preset execution speed threshold. If the execution speed is less than the preset execution speed threshold, the corresponding processing component in the multimodal component set is updated to increase the execution speed of the AI ​​content generation thread.

9. The multimodal AIGC generation service control system as described in claim 8, characterized in that, The multimodal AIGC generation service control system also includes a module for performing the following operations: If the execution speed is not less than a preset execution speed threshold, the first running parameters of the AI ​​content generation thread are monitored in real time. The first running parameters include the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate. The first operation evaluation coefficient is obtained by using the amount of content data generated per unit time, memory utilization, the difference in data volume between input content and output content, and the content generation error rate included in the first operation parameters. The first operational evaluation coefficient is obtained using the following formula: Among them, R 01 denoted as the first performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi M represents the average difference in data volume between the input and output data corresponding to the i-th unit of time; i E represents the memory utilization rate corresponding to the i-th unit of time; i This represents the bit error rate of content generation corresponding to the i-th unit of time; The first performance evaluation coefficient is compared with a preset coefficient threshold. When the first performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation and an abnormal risk warning is issued. If the execution speed is less than a preset execution speed threshold, the second running parameters of the AI ​​content generation thread are monitored in real time. The second running parameters include the amount of content data generated per unit time, the difference in data volume between the input content and the output content, and the execution speed value. The second operation evaluation coefficient is obtained by using the amount of content data generated per unit time, the difference in data volume between input content and output content, and the execution speed value contained in the second operation parameters. The second operational evaluation coefficient is obtained using the following formula: Among them, R 02 denoted as the second performance evaluation coefficient; n represents the number of time units experienced by the AI ​​content generation thread; C i C represents the amount of content data generated in the i-th unit of time; zpi C represents the average difference between the input and output data at the i-th unit of time; zmaxi This represents the maximum difference in data volume between the input and output data corresponding to the i-th unit of time; v i This represents the execution speed value generated in the i-th unit of time; v b This represents the standard deviation of the generated execution speed value corresponding to n units of time; The second performance evaluation coefficient is compared with a preset coefficient threshold. When the second performance evaluation coefficient exceeds the preset coefficient threshold, it is determined that the AI ​​content generation thread has a risk of abnormal operation, and an abnormal risk warning is issued.

10. The multimodal AIGC generation service control system as described in claim 6, characterized in that: The AI ​​content generation task judgment module is used to detect the final generated content of the AI ​​content generation task and determine whether the AI ​​content generation task has been executed correctly, including: The final generated content of the AI ​​content generation task is detected to obtain the data error rate of the final generated content; if the data error rate is greater than or equal to a preset error rate threshold, it is determined that the AI ​​content generation task has not been executed correctly; otherwise, it is determined that the AI ​​content generation task has been executed correctly. The multimodal component set operation module is used to output the final generated content when the AI ​​content generation task is executed correctly, and to integrate and save the multimodal component set currently loaded on the user terminal, including: When the AI ​​content generation task is executed correctly, the final generated content is backed up and then output. The set of multimodal components currently loaded on the user terminal is integrated into a set of reusable multimodal components and stored in a designated storage space within the user terminal.

Citation Information

Patent Citations

  • Video generation method and device, live broadcast processing method and device and readable medium

    CN113923462A

  • Video synthesis method and electronic equipment

    CN117793271A