Audio and video processing method and system based on artificial intelligence

By building a matching tree in the audio and video processing platform and deploying the processing model to the front-end device of the user, the problem of cloud audio and video processing delay and resource competition is solved, and efficient audio and video processing and optimized user experience is achieved.

CN120075481AInactive Publication Date: 2025-05-30BEIJING JINGAN XINDA TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510362654.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, audio and video processing models are usually deployed in the cloud, resulting in a long delay in uploading and processing of audio and video data and resource competition.

Method used

By building an audio and video processing platform, it receives audio and video tasks uploaded by registrants, divides active areas, selects intended users and front-end devices, creates a collection of processing models, and deploys the processing models to the front-end devices through a matching tree to realize distributed computing.

Benefits of technology

Maximize the use of computing resources, reduce redundant computing and resource waste, optimize computing resource allocation, reduce network latency, alleviate cloud pressure, and improve the processing efficiency and user experience of audio and video tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075481A_ABST
    Figure CN120075481A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of audio and video processing, and particularly relates to an audio and video processing method and system based on artificial intelligence, and the method comprises the steps: constructing an audio and video processing platform, receiving an audio and video task uploaded by a registrant, dividing an active region of the audio and video task, selecting an intentional user, and defining a front-end device. Reading attribute data, wherein the attribute data at least comprises bandwidth, CPU (Central Processing Unit) performance and storage space; the method comprises the following steps: creating a processing model set of audio and video tasks, configuring a calculation demand of each processing model, creating left nodes in one-to-one correspondence with the processing models, and uploading the calculation demands to the left nodes. By constructing the matching tree, the processing model can be deployed to different front-end devices, so that the network delay is greatly reduced, the cloud pressure is greatly relieved through distributed calculation of calculation requirements, and the processing efficiency of audio and video tasks and the user experience are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio - video processing, and in particular, to an audio - video processing method and system based on artificial intelligence. Background Art

[0002] When using an artificial intelligence model to process audio - video data, a large amount of data consumption often occurs. The artificial intelligence model is generally deployed in the cloud. Although the powerful computing power of the cloud can be used to complete complex analysis tasks, this method also has certain limitations. First, due to the large volume of audio - video data, the process of uploading it to the cloud may take a long time, especially when the network bandwidth is limited. Second, when a large number of devices and users send data to the cloud simultaneously, the cloud server may face problems such as tight resource allocation and piled - up computing tasks, which will further exacerbate data congestion and processing delay.

[0003] Therefore, "how to deploy the audio - video processing model on the user side" is the technical problem to be solved by the present invention. Summary of the Invention

[0004] The purpose of the present invention is to provide an audio - video processing method and system based on artificial intelligence to solve the problem of "how to deploy the audio - video processing model on the user side" proposed in the above - mentioned background art.

[0005] To achieve the above - mentioned purpose, the present invention provides the following technical solutions:

[0006] An audio - video processing method based on artificial intelligence, the method includes:

[0007] Construct an audio - video processing platform, receive the audio - video tasks uploaded by registrants, divide the active area of the audio - video tasks, select intended users, define front - end devices, and read out attribute data, where the attribute data includes at least: bandwidth, CPU performance, and storage space;

[0008] Create a set of processing models for the audio - video tasks, configure the computing requirements for each processing model, create left nodes corresponding to the processing models one by one, upload the computing requirements to the left nodes, and sort the left nodes in descending order of the computing requirements to obtain a queue;

[0009] Create right nodes corresponding to the front - end devices one by one, and through the right nodes, traverse in order from the head to the tail of the queue to find the front - end device that first meets the computing requirements of the left node, and establish the corresponding relationship between the left node and the right node, and integrate the corresponding relationship to generate a matching tree;

[0010] Obtain the call permission of the front-end device via the matching tree, deploy the corresponding processing model to the front-end device, select a target model based on the audio-visual task, define the front-end device corresponding to the target model as the target device, and use the target device to process the audio-visual task.

[0011] Further, the steps of constructing the audio-visual processing platform and receiving the audio-visual tasks uploaded by the registrants and dividing the active area of the audio-visual tasks include:

[0012] Trace back the source point of the audio-visual task and mark it on the preset geographical information map to generate a task heat map;

[0013] Define the area boundary, generate several task areas, calculate the number of audio-visual tasks in each task area, and define the task areas with the number greater than the threshold as the active areas.

[0014] Further, the steps of dividing the active area of the audio-visual task, selecting the intended users, and defining the front-end device include:

[0015] Embed a registration mechanism into the audio-visual processing platform and configure the roles of each registrant, where the roles at least include: ordinary users and node providers;

[0016] Define the node providers located in the active area as the intended users, define the terminal devices corresponding to the intended users as the front-end devices, and obtain the usage permissions of the front-end devices.

[0017] Further, the steps of creating a set of processing models for the audio-visual task and configuring the computing requirements of each processing model include:

[0018] Configure the influencing factors of the computing requirements, where the influencing factors at least include: the size and complexity of the data input, and dynamically adjust the computing requirements;

[0019] In the active area, select standby nodes, set the height of the queue, and activate the standby nodes when the left nodes exceed the height.

[0020] Further, the steps of finding the front-end device that first meets the computing requirements of the left node, establishing the corresponding relationship between the left node and the right node, and integrating the corresponding relationship to generate a matching tree include:

[0021] Integrate the root node into the queue and establish the overall relationship between the root node and the left node and the right node;

[0022] Embed a tilting mechanism into the matching tree, where the tilting mechanism is: when the attribute data in the right node cannot meet the computing requirements in the left node, tilt the matching tree.

[0023] Further, the step of defining the front-end device corresponding to the target model as the target device and using the target device to process the audio-video task includes:

[0024] Configure the priority of the audio-video task and construct a priority handling strategy;

[0025] Establish a communication link between the target device and the source point, and adjust the communication link using the priority handling strategy.

[0026] Further, the method further includes:

[0027] Configure the usage scenario of the audio-video task, and insert preference tags into the front-end device;

[0028] Establish a mapping between the chunks and the front-end device via the preference tags.

[0029] Further, the system includes:

[0030] A reading module, configured to construct an audio-video processing platform, receive the audio-video tasks uploaded by registrants, divide the active area of the audio-video tasks, select the intended users, define the front-end devices, and read out the attribute data, where the attribute data includes at least: bandwidth, CPU performance, and storage space;

[0031] A obtaining module, configured to create a set of processing models for the audio-video tasks, configure the computing requirements of each processing model, create left nodes corresponding to the processing models one by one, upload the computing requirements to the left nodes, and sort the left nodes in descending order of the computing requirements to obtain a queue;

[0032] A generating module, configured to create right nodes corresponding to the front-end devices one by one, and through the right nodes, traverse in sequence from the head to the tail of the queue, find the front-end device that first meets the computing requirements of the left node, and establish the corresponding relationship between the left node and the right node, and integrate the corresponding relationship to generate a matching tree;

[0033] A processing module, configured to obtain the call permission of the front-end device via the matching tree, deploy the corresponding processing model to the front-end device, select the target model based on the audio-video task, define the front-end device corresponding to the target model as the target device, and use the target device to process the audio-video task.

[0034] Further, the reading module includes:

[0035] A backtracking unit, configured to backtrack the source point of the audio-video task and mark it on a preset geographical information map to generate a task heat map;

[0036] A defining unit, configured to delimit a region boundary, generate a plurality of task areas, calculate the number of audio-visual tasks in each task area, and define a task area with the number greater than a threshold as an active area;

[0037] A configuration unit, configured to embed a registration mechanism into an audio-visual processing platform and configure the roles of each registrant, where the roles at least include: ordinary users and node providers;

[0038] An obtaining unit, configured to define a node provider located in the active area as an intended user, define a terminal device corresponding to the intended user as a front-end device, and obtain the usage permission of the front-end device.

[0039] Further, the obtaining module includes:

[0040] An adjustment unit, configured to configure influencing factors of the computing requirement, where the influencing factors at least include: the size and complexity of data input, and dynamically adjust the computing requirement;

[0041] An activation unit, configured to select standby nodes in the active area, set the height of a queue, and activate the standby nodes when a left node exceeds the height.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] By constructing an audio-visual processing platform, it can adapt to different audio-visual scenarios and task requirements, maximize the utilization of computing resources, avoid redundant calculations and resource waste caused by parallel operation of multiple models. By determining the computing requirement of a processing model, it can optimize the allocation of computing resources, avoid performance bottlenecks or task delays caused by resource contention, and at the same time can also achieve dynamic scheduling of computing resources, flexibly arrange the execution order of processing models, prevent overloading or resource shortage. By constructing a matching tree, it can deploy the processing model to different front-end devices, thereby greatly reducing network latency, and through distributed computing of the computing requirement, greatly alleviating the cloud pressure and greatly improving the processing efficiency of audio-visual tasks and user experience. Description of the Drawings

[0044] Figure 1 It is a flowchart of an audio-visual processing method based on artificial intelligence provided by an embodiment of the present invention;

[0045] Figure 2 It is a first sub-flowchart of an audio-visual processing method based on artificial intelligence provided by an embodiment of the present invention;

[0046] Figure 3It is the second sub - process block diagram of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention;

[0047] Figure 4 It is the third sub - process block diagram of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention;

[0048] Figure 5 It is the fourth sub - process block diagram of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention;

[0049] Figure 6 It is the block diagram of the composition of the audio - video processing system based on artificial intelligence provided by the embodiments of the present invention;

[0050] Figure 7 It is the block diagram of the composition of the reading module in the audio - video processing system based on artificial intelligence provided by the embodiments of the present invention;

[0051] Figure 8 It is the block diagram of the composition of the obtaining module in the audio - video processing system based on artificial intelligence provided by the embodiments of the present invention;

[0052] Figure 9 It is the block diagram of the composition of the generating module in the audio - video processing system based on artificial intelligence provided by the embodiments of the present invention;

[0053] Figure 10 It is the block diagram of the composition of the processing module in the audio - video processing system based on artificial intelligence provided by the embodiments of the present invention. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] In Embodiment 1, Figure 1 The implementation process of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention is shown, and the details are as follows:

[0056] S100: Construct an audio - video processing platform, receive the audio - video tasks uploaded by registrants, divide the active areas of the audio - video tasks, select the intended users, define the front - end devices, and read out the attribute data, where the attribute data at least includes: bandwidth, CPU performance, and storage space.

[0057] Build an audio - video processing platform, open user registration, and identify the registered users. The audio - video processing platform is mainly used to process the audio - video tasks uploaded by the registered users. The specific processing methods include: synchronization, alignment, format conversion, semantic recognition, and adding subtitles, etc.; Trace the location where each audio - video task is initiated, and divide the active area based on this location. The active area refers to the area where audio - video tasks need to be frequently processed; For example, the active area can be the gathering place of online content creators or a certain e - commerce park, etc.; Select from the registered users the part that is willing to provide terminal devices for the deployment of the processing model, that is, the intended users; Define the terminal devices corresponding to the intended users as front - end devices. The front - end devices can be personal computers or servers, etc. Determine the attribute data of each front - end device. The attribute data includes: the bandwidth of the device, CPU performance, and storage space, etc. The bandwidth can reflect the network transmission ability of the front - end device and determine the speed of audio - video data upload, download, and real - time stream processing. The CPU performance represents the computing ability of the front - end device and directly affects the execution efficiency of computing tasks such as audio - video decoding, encoding, and audio - video processing.

[0058] S200: Create a set of processing models for audio - video tasks, configure the computing requirements for each processing model, create left nodes corresponding to the processing models one by one, upload the computing requirements to the left nodes, and sort the left nodes in descending order of the computing requirements to obtain a queue.

[0059] Create a set of processing models composed of several processing models. The processing models include: audio noise reduction model, speech recognition model, video segmentation model, object detection model, and image quality enhancement model, etc.; Determine the computing requirements for each processing model, clarify the key resource indicators required for the processing model during actual operation, including CPU computing power, memory occupancy, bandwidth requirements, and storage space, etc. The computing requirements are also the hardware requirements when deploying the processing model; For example, the speech recognition model may require high CPU computing power and stable bandwidth to ensure the smoothness of real - time transcription; The video super - resolution model may require more memory and storage space to support the frame - by - frame processing of high - definition images; Through these configuration data, the audio - video processing platform can not only allocate the most suitable model for different types of front - end devices, but also dynamically adjust resource allocation during task scheduling; Create several left nodes. The left nodes are mainly used to store processing models, and each left node can only store one processing model. Upload the computing requirements corresponding to each processing model to the corresponding left nodes, and sort all the left nodes in descending order of the computing requirements to obtain a queue. The queue is a set composed of left nodes.

[0060] S300: Create right nodes corresponding one by one to the front-end devices. Via the right nodes, traverse them in sequence according to the order from the head to the tail of the queue, find out the front-end device that first meets the computing requirements of the left node, and establish the corresponding relationship between the left node and the right node, and integrate the corresponding relationship to generate a matching tree.

[0061] Create right nodes corresponding to each front-end device. The right node is a virtual mapping unit of the front-end device, which is used to record the attribute data, processing capabilities, and task allocation information of the front-end device. Use each right node to compare with the left nodes in the queue in turn. The comparison order is from the head to the tail of the queue. Find out the right node that can first meet the computing requirements in the left node, and establish the corresponding relationship. Mount the left node and the right node with the corresponding relationship to the same parent node, and integrate the parent node and the corresponding left node and right node to generate a matching tree. Among them, the matching tree can reconstruct the corresponding relationship, which is similar to the binary tree in the prior art.

[0062] For example, there are four groups of left nodes, namely A, B, C, and D, and they form a queue in the order of B - C - A - D. In other words, the processing model corresponding to the head of the queue has the highest requirement for computing power, that is, the processing model in the B left node has the highest requirement for computing power, and the processing model in the D left node has the lowest requirement for computing power. There is a right node X. According to the data such as its corresponding bandwidth, CPU performance, and storage space, compare it with the left nodes in the queue in turn, and find out the left node that can be satisfied first. If the right node X cannot meet the deployment requirements of the B processing model, but meets the deployment requirements of the three processing models C, A, and D, then establish the corresponding relationship between X and C, and mount the two to the same parent node, and integrate the parent node, the left node C, and the right node X to generate a matching tree.

[0063] S400: Via the matching tree, obtain the call permission of the front-end device, deploy the corresponding processing model to the front-end device, select the target model based on the audio-visual task, define the front-end device corresponding to the target model as the target device, and use the target device to process the audio-visual task.

[0064] According to the audio-visual task, determine the processing model to be used, that is, the target model, and the front-end device where the target model is located is the target device. According to the call permission, send the audio-visual task to the target device, and use the processing model in the target device to process the audio-visual task.

[0065] In Embodiment 2, Figure 2 It shows the implementation process of the audio-visual processing method based on artificial intelligence provided by the embodiment of the present invention. The steps of constructing the audio-visual processing platform, receiving the audio-visual task uploaded by the registrant, and dividing the active area of the audio-visual task are described in detail as follows:

[0066] S101: Trace back the source point of the audio - video task and mark it on a preset geographical information map to generate a task heat map.

[0067] The audio - video processing platform traces back the source point of the task through log records and task identifiers, that is, which user, device or region initiated the task, and marks the source point on the geographical information map to generate a task heat map; the task heat map can intuitively display the distribution of the source points, and determine the density and activity of audio - video tasks at different geographical locations.

[0068] S102: Define the regional boundaries to generate several task areas, calculate the number of audio - video tasks in each task area, and define the task areas with the number greater than the threshold as active areas.

[0069] In the task heat map, divide into multiple task areas, calculate the number of audio - video tasks in each task area, and if the number is greater than the threshold, define the corresponding task area as an active area.

[0070] In Embodiment 3, Figure 2 The implementation process of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention is shown. The following details the steps of dividing the active areas of audio - video tasks, selecting intended users, and defining front - end devices, as follows:

[0071] S103: Embed a registration mechanism into the audio - video processing platform and configure the roles of each registrant, where the roles at least include: ordinary users and node providers.

[0072] Embed a registration mechanism into the audio - video processing platform. The registration mechanism is as follows: allow users to register as platform members (i.e., registrants) through personal information, account binding, device authorization, etc.; during the registration process, configure the roles of each registrant according to the different needs and responsibilities of users to ensure that users with different roles have different access rights in the platform; the roles at least include two types: ordinary users and node providers, where node providers are ordinary users who are willing to provide terminal devices for deploying processing models.

[0073] S104: Define the node providers located in the active area as intended users, define the terminal devices corresponding to the intended users as front - end devices, and obtain the usage rights of the front - end devices.

[0074] Use the usage rights to deploy the processing model into the front - end devices of the intended users.

[0075] In Embodiment 4, Figure 3The implementation process of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention is shown. The following details the steps of creating a set of processing models for the audio - video task and configuring the computing requirements for each processing model, as follows:

[0076] S201: Configure the influencing factors of the computing requirements, where the influencing factors at least include: the size and complexity of the data input, and dynamically adjust the computing requirements.

[0077] When determining the computing requirements for each processing model, in addition to considering the complexity of the processing model, it is also necessary to consider the size and complexity of the data input, that is, the influencing factors, and dynamically adjust the computing requirements according to the influencing factors.

[0078] S202: In the active area, select standby nodes, set the height of the queue, and activate the standby nodes when the left nodes exceed the height.

[0079] In the active area, select standby nodes, where the standby nodes are also standby front - end devices. Set a height for each left node and set a height for the queue. The queue height is determined by the number of front - end devices. If the sum of the heights of the left nodes is greater than the height of the queue, activate the standby nodes.

[0080] For example, if the height of the left node is set to 1 centimeter and there are 10 front - end devices available for deploying the processing model, the height of the queue is 10 centimeters. However, if there are 12 processing models to be deployed, at this time, activate the standby nodes and deploy the extra 2 processing models to the standby nodes.

[0081] In Embodiment 5, Figure 4 The implementation process of the audio - video processing method based on artificial intelligence provided by the embodiments of the present invention is shown. The following details the steps of finding the front - end device that first meets the computing requirements of the left node, establishing the corresponding relationship between the left node and the right node, and integrating the corresponding relationships to generate a matching tree, as follows:

[0082] S301: Integrate the root node into the queue and establish the overall relationship between the root node, the left node, and the right node.

[0083] In the queue, create a root node, which is mainly used to dynamically adjust the corresponding relationship between the left node and the right node.

[0084] Continue to detail the example in S300. If the processing volume of the audio - video task is large, resulting in the inability of the left node C to be processed quickly, then find another device (assumed to be Y) with better performance than X from the front - end devices, and then re - establish the corresponding relationship between the left node C and the right node Y.

[0085] S302: Embed a tilting mechanism into the matching tree, where the tilting mechanism is: when the attribute data in the right node cannot meet the calculation requirements in the left node, tilt the matching tree.

[0086] In the above, if the left node C cannot be processed quickly, tilt the matching tree composed of the left node C and the right node X, so as to intuitively display the processing progress of the audio-visual task.

[0087] In Embodiment 6, Figure 5 The implementation process of the audio-visual processing method based on artificial intelligence provided by the embodiments of the present invention is shown. The following details the steps of defining the front-end device corresponding to the target model as the target device and using the target device to process the audio-visual task, as follows:

[0088] S401: Configure the priority of the audio-visual task and construct a priority handling strategy.

[0089] Determine the priority of the audio-visual task according to the urgency of the audio-visual task. The priorities are divided into high, medium, and low, and the allocation of priorities is determined by the operation and maintenance personnel of the audio-visual processing platform. The priority handling strategy is: preferentially process the audio-visual tasks corresponding to high priorities.

[0090] S402: Establish a communication link between the target device and the source point, and adjust the communication link using the priority handling strategy.

[0091] After determining the target device, establish a communication link between the target user and the source point, and preferentially connect the communication link to the front-end device corresponding to the high priority.

[0092] In Embodiment 7, different from Embodiment 1, in the embodiments of the present invention, the method further includes:

[0093] Configure the usage scenario of the audio-visual task and insert preference tags into the front-end device;

[0094] Establish a mapping between the blocks and the front-end device via the preference tags.

[0095] Determine the usage scenarios for each audio-visual task. The usage scenarios can be live broadcast scenarios, video editing scenarios, video surveillance scenarios, etc. When the number of front-end devices is relatively sufficient, the processing model can be adjusted according to the performance, type, and characteristics of the audio-visual tasks to be processed. After the adjustment, preference tags are inserted into the corresponding front-end devices. For example, if there are multiple front-end devices with an audio noise reduction model deployed, the audio noise reduction model can be adjusted to meet the usage requirements in different scenarios, that is, each scenario corresponds to an audio noise reduction model, and the audio noise reduction models in different scenarios are deployed to different front-end devices.

[0096] Figure 6 The block diagram showing the composition structure of the audio-visual processing system based on artificial intelligence provided by the embodiment of the present invention, the audio-visual processing system 1 based on artificial intelligence includes:

[0097] A reading module 11, configured to build an audio-visual processing platform, receive the audio-visual tasks uploaded by registrants, divide the active area of the audio-visual tasks, select the intended users, define the front-end devices, and read out the attribute data, where the attribute data at least includes: bandwidth, CPU performance, and storage space;

[0098] A obtaining module 12, configured to create a set of processing models for the audio-visual tasks, configure the computing requirements for each processing model, create left nodes corresponding to the processing models one by one, upload the computing requirements to the left nodes, and sort the left nodes in descending order of the computing requirements to obtain a queue;

[0099] A generating module 13, configured to create right nodes corresponding to the front-end devices one by one, and through the right nodes, traverse in sequence from the head to the tail of the queue to find the front-end device that first meets the computing requirements of the left node, and establish the corresponding relationship between the left node and the right node, and integrate the corresponding relationship to generate a matching tree;

[0100] A processing module 14, configured to obtain the call permission of the front-end device through the matching tree, deploy the corresponding processing model to the front-end device, select a target model based on the audio-visual task, define the front-end device corresponding to the target model as the target device, and use the target device to process the audio-visual task.

[0101] Figure 7 The block diagram showing the composition structure of the audio-visual processing system based on artificial intelligence provided by the embodiment of the present invention, the reading module 11 includes:

[0102] A backtracking unit 111, configured to backtrack the source point of the audio-visual task and mark it on a preset geographical information map to generate a task heat map;

[0103] Define unit 112, which is used to delimit the area boundary, generate a number of task areas, calculate the number of audio-visual tasks in each task area, and define the task areas with the number greater than the threshold as active areas;

[0104] Configure unit 113, which is used to embed a registration mechanism into the audio-visual processing platform and configure the roles of each registrant, where the roles at least include: ordinary users and node providers;

[0105] Obtain unit 114, which is used to define the node providers located in the active area as intended users, define the terminal devices corresponding to the intended users as front-end devices, and obtain the usage permissions of the front-end devices.

[0106] Figure 8 The block diagram of the composition structure of the audio-visual processing system based on artificial intelligence provided by the embodiment of the present invention is shown. The obtaining module 12 includes:

[0107] Adjust unit 121, which is used to configure the influencing factors of the computing requirements, where the influencing factors at least include: the size and complexity of data input, and dynamically adjust the computing requirements;

[0108] Activate unit 122, which is used to select standby nodes in the active area, set the height of the queue, and activate the standby nodes when the left node exceeds the height.

[0109] Figure 9 The block diagram of the composition structure of the audio-visual processing system based on artificial intelligence provided by the embodiment of the present invention is shown. The generating module 13 includes:

[0110] Establish unit 131, which is used to integrate root nodes into the queue and establish the overall relationship between the root nodes and the left and right nodes;

[0111] Tilt unit 132, which is used to embed a tilt mechanism into the matching tree, where the tilt mechanism is: when the attribute data in the right node cannot meet the computing requirements in the left node, tilt the matching tree.

[0112] Figure 10 The block diagram of the composition structure of the audio-visual processing system based on artificial intelligence provided by the embodiment of the present invention is shown. The processing module 14 includes:

[0113] Construct unit 141, which is used to configure the priority of the audio-visual tasks and construct a priority handling strategy;

[0114] Build unit 142, which is used to build a communication link between the target device and the source point and adjust the communication link by using the priority handling strategy.

[0115] Among them, the reading module 11 is mainly used to complete step S100, the obtaining module 12 is mainly used to complete step S200, the generating module 13 is mainly used to complete step S300, and the processing module 14 is mainly used to complete step S400;

[0116] The backtracking unit 111 is mainly used to complete step S101, the defining unit 112 is mainly used to complete step S102, the configuring unit 113 is mainly used to complete step S103, and the obtaining unit 114 is mainly used to complete step S104;

[0117] The adjusting unit 121 is mainly used to complete step S201, and the activating unit 122 is mainly used to complete step S202;

[0118] The establishing unit 131 is mainly used to complete step S301, and the tilting unit 132 is mainly used to complete step S302;

[0119] The constructing unit 141 is mainly used to complete step S401, and the building unit 142 is mainly used to complete step S402.

[0120] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0121] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

[0122] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An audio and video processing method based on artificial intelligence, characterized in that: The method comprises: Build an audio and video processing platform, receive audio and video tasks uploaded by registrants, divide active areas of audio and video tasks, select intended users, define front-end devices, and read attribute data, where the attribute data at least includes: bandwidth, CPU performance and storage space; Create a set of processing models for audio and video tasks, configure the computing requirements of each processing model, create left nodes corresponding to the processing models one by one, upload the computing requirements to the left nodes, and sort the left nodes in descending order of computing requirements to obtain a queue; Create a right node corresponding to the front-end device one by one, traverse the queue from the head to the tail through the right node, find the front-end device that first meets the computing requirements of the left node, establish a corresponding relationship between the left node and the right node, integrate the corresponding relationship, and generate a matching tree; Through the matching tree, the calling authority of the front-end device is obtained, the corresponding processing model is deployed to the front-end device, based on the audio and video task, the target model is selected, the front-end device corresponding to the target model is defined as the target device, and the target device is used to process the audio and video task.

2. The audio and video processing method based on artificial intelligence according to claim 1 is characterized in that: The steps of constructing an audio and video processing platform, receiving audio and video tasks uploaded by registrants, and dividing active areas of audio and video tasks include: Tracing back the source point of the audio and video task, marking it in a preset geographic information map, and generating a task heat map; The region boundary is delineated, a number of task areas are generated, the number of audio and video tasks in each task area is calculated, and the task area with the number greater than a threshold is defined as an active area.

3. The audio and video processing method based on artificial intelligence according to claim 1 is characterized in that: The steps of dividing the active area of ​​the audio and video task, selecting the intended user, and defining the front-end device include: Embed a registration mechanism into the audio and video processing platform and configure the role of each registrant, wherein the roles include at least: ordinary user and node provider; A node provider located in the active area is defined as an intended user, and a terminal device corresponding to the intended user is defined as a front-end device, and the use authority of the front-end device is obtained.

4. The audio and video processing method based on artificial intelligence according to claim 1 is characterized in that: The steps of creating a set of processing models for audio and video tasks and configuring the computing requirements of each processing model include: Configuring the influencing factors of the computing demand, wherein the influencing factors include at least: the size and complexity of the data input, and dynamically adjusting the computing demand; In the active area, a standby node is selected, and a queue height is set. When the left node exceeds the height, the standby node is activated.

5. The audio and video processing method based on artificial intelligence according to claim 4 is characterized in that: The steps of finding out the front-end device that first meets the computing requirements of the left node, establishing a corresponding relationship between the left node and the right node, integrating the corresponding relationship, and generating a matching tree include: Integrate the root node into the queue, and establish a coordinated relationship between the root node and the left node and the right node; A tilting mechanism is embedded in the matching tree, wherein the tilting mechanism is: when the attribute data in the right node cannot meet the calculation requirements in the left node, the matching tree is tilted.

6. The audio and video processing method based on artificial intelligence according to claim 2 is characterized in that: The step of defining the front-end device corresponding to the target model as a target device and using the target device to process the audio and video task includes: Configure the priority of the audio and video tasks and build a priority handling strategy; A communication link is established between the target device and the source point, and the communication link is adjusted using the priority handling strategy.

7. The audio and video processing method based on artificial intelligence according to claim 1 is characterized in that: The method further comprises: Configuring a usage scenario for the audio and video task and inserting a preference tag into the front-end device; A mapping between the blocks and the front-end devices is established via the preference tags.

8. An audio and video processing system based on artificial intelligence, characterized in that: The system comprises: The reading module is used to build an audio and video processing platform, receive audio and video tasks uploaded by registrants, divide active areas of audio and video tasks, select intended users, define front-end devices, and read attribute data, wherein the attribute data at least includes: bandwidth, CPU performance and storage space; A module is obtained, which is used to create a set of processing models for audio and video tasks, configure the computing requirements of each processing model, create left nodes corresponding to the processing models one by one, upload the computing requirements to the left nodes, and sort the left nodes in descending order of computing requirements to obtain a queue; A generation module is used to create a right node corresponding to the front-end device one by one, and traverse the queue from the head to the tail in sequence through the right node to find the front-end device that first meets the computing requirements of the left node, establish a corresponding relationship between the left node and the right node, integrate the corresponding relationship, and generate a matching tree; The processing module is used to obtain the calling authority of the front-end device through the matching tree, deploy the corresponding processing model to the front-end device, select the target model based on the audio and video task, define the front-end device corresponding to the target model as the target device, and use the target device to process the audio and video task.

9. The audio and video processing system based on artificial intelligence according to claim 8, characterized in that: The reading module comprises: A tracing unit, used to trace back the source point of the audio and video task, mark it in a preset geographic information map, and generate a task heat map; A definition unit, used to define the region boundary, generate a plurality of task areas, calculate the number of audio and video tasks in each task area, and define the task area with the number greater than a threshold as an active area; A configuration unit, used to embed a registration mechanism into the audio and video processing platform, and configure a role for each registrant, wherein the roles include at least: a common user and a node provider; The acquisition unit is used to define the node provider located in the active area as the intended user, define the terminal device corresponding to the intended user as the front-end device, and obtain the use authority of the front-end device.

10. The audio and video processing system based on artificial intelligence according to claim 8, characterized in that: The obtaining module comprises: An adjustment unit, configured to configure influencing factors of the computing demand, wherein the influencing factors include at least: the size and complexity of data input, and dynamically adjust the computing demand; The activation unit is used to select a standby node in the active area, set the height of the queue, and activate the standby node when the left node exceeds the height.

Citation Information

Cited By

  • Information data acquisition method and system based on artificial intelligence

    CN120564233A