A cloud-edge collaboration-based multi-modal analysis method, device, equipment and medium

By allocating analysis process nodes to cloud terminals and edge terminals in the cloud-edge collaborative system, the resource waste and scheduling problems of multimodal solutions in cloud-edge scenarios are solved, and efficient multimodal data analysis is achieved.

CN114613019BActive Publication Date: 2025-10-28ASIAINFO TECH CHINA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111602778.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2025-10-28
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

Existing multimodal solutions lack support in cloud-edge scenarios, resulting in wasted computing power and improper resource scheduling, and there are few mature models available.

Method used

By assigning analysis process nodes of analysis tasks to cloud terminals and edge terminals, and performing analysis operations when it is determined that the current node is not configured with a duplicate identifier, the resources of cloud/edge terminals are used to collaboratively process multimodal data, thereby achieving cloud-edge collaborative analysis.

Benefits of technology

It enables efficient processing of multimodal data analysis tasks in a cloud-edge collaborative environment, avoiding resource waste and performance issues, and supports multimodal analysis in cloud-edge scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114613019B_ABST
    Figure CN114613019B_ABST
Patent Text Reader

Abstract

This application provides a cloud-edge collaborative multimodal analysis method, apparatus, electronic device, computer-readable storage medium, and computer program product, relating to the field of artificial intelligence. The method is applied to cloud / edge terminals. The method applied to cloud terminals includes: allocating analysis nodes of the analysis process corresponding to the analysis task to both the cloud terminal and the edge terminal; during the execution of the analysis process, after determining that the current analysis node is located on the cloud terminal, if the current analysis node is not configured with a duplicate identifier, performing corresponding analysis operations, including analyzing multimodal data, to obtain analysis results; then, after determining that the next analysis node is located on the edge terminal, sending the analysis results to the edge terminal; otherwise, continuing to execute the analysis operations of the next analysis node on the cloud terminal. This method achieves the purpose of cloud-edge collaborative processing by comprehensively utilizing the resources of cloud / edge terminals to process multimodal data analysis tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a multimodal analysis method, apparatus, electronic device, computer-readable storage medium, and computer program product based on cloud-edge collaboration. Background Technology

[0002] Currently, existing single-modal, multimodal, and orchestration-based solutions can all be used as artificial intelligence analysis solutions to address the corresponding problems.

[0003] Specifically, among the three solutions mentioned above, the unimodal solution is more commonly used. However, the unimodal solution can only solve problems in simple, single scenarios, and its recognition accuracy is low or it cannot recognize complex, personalized scenarios. Therefore, multimodal solutions have emerged, mainly focusing on multimodal recognition of speech and vision, but mature models are few and far between. Consequently, a solution based on process orchestration has been developed to implement multimodal recognition. However, this solution has relatively simple orchestration and analysis capabilities, suffers from wasted computing power due to repeated execution, performance issues due to a lack of consideration for resource scheduling, and does not support cloud-edge scenarios.

[0004] Therefore, the existing multimodal solutions do not offer significant advantages and do not support cloud-edge scenarios. Summary of the Invention

[0005] The purpose of this application is to address the lack of support for multimodal solutions in cloud-edge scenarios.

[0006] According to a first aspect of the embodiments of this application, a multimodal analysis method based on cloud-edge collaboration is provided, applied to a cloud terminal, the method comprising:

[0007] Assign analysis nodes of the analysis process corresponding to the analysis task to cloud terminals and edge terminals;

[0008] After determining that the current analysis node in the analysis process is located on a cloud terminal, if the current analysis node is not configured with a duplicate identifier, the analysis operation corresponding to the current analysis node is executed to obtain the analysis results. The analysis operation includes analyzing the multimodal data to be analyzed using an AI model.

[0009] If the next analysis node of the current analysis node is at the edge terminal, the analysis results will be sent to the edge terminal.

[0010] According to a second aspect of the embodiments of this application, a multimodal analysis method based on cloud-edge collaboration is provided, applied to an edge terminal, the method comprising:

[0011] Receive the analysis nodes of the analysis process corresponding to the analysis task assigned by the cloud terminal;

[0012] After determining that the current analysis node in the analysis process is at the edge terminal, if the current analysis node is not configured with a duplicate identifier, the analysis operation corresponding to the current analysis node is executed to obtain the analysis result. The analysis operation includes at least analyzing the multimodal data to be analyzed using an AI model.

[0013] If the next analysis node of the current analysis node is in the cloud terminal, the analysis results will be sent to the cloud terminal.

[0014] According to a third aspect of the embodiments of this application, a multimodal analysis device based on cloud-edge collaboration is provided, applied to a cloud terminal, the device comprising:

[0015] The allocation module is used to allocate the analysis nodes of the analysis process corresponding to the analysis task to cloud terminals and edge terminals;

[0016] The execution module is used to perform the analysis operation corresponding to the current analysis node after determining that the current analysis node of the analysis process is in the cloud terminal. If the current analysis node is not configured with a duplicate identifier, the analysis operation is performed to obtain the analysis result. The analysis operation includes analyzing the multimodal data to be analyzed through an AI model.

[0017] The transceiver module is used to send the analysis results to the edge terminal if the next analysis node of the current analysis node is located at the edge terminal.

[0018] According to a fourth aspect of the embodiments of this application, a multimodal analysis device based on cloud-edge collaboration is provided, applied to an edge terminal, the device comprising:

[0019] The transceiver module is used to receive analysis nodes of the analysis process corresponding to the analysis task allocated by the cloud terminal;

[0020] The execution module is used to perform the analysis operation corresponding to the current analysis node after determining that the current analysis node in the analysis process is at the edge terminal. If the current analysis node is not configured with a duplicate identifier, the analysis operation will be performed to obtain the analysis result. The analysis operation includes at least analyzing the multimodal data to be analyzed through an AI model.

[0021] The transceiver module is used to send the analysis results to the cloud terminal if the next analysis node of the current analysis node is located in the cloud terminal.

[0022] According to a fifth aspect of the embodiments of this application, an electronic device is provided, the electronic device including: a memory, a processor and a computer program stored in the memory, the processor executing the computer program to implement the steps of the method shown in either the first or second aspect of this application.

[0023] According to a sixth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method shown in either the first or second aspect of this application.

[0024] According to one aspect of the embodiments of this application, a computer program product is provided, the product including a computer program that, when executed by a processor, implements the steps of the method shown in either the first or second aspect of this application.

[0025] The beneficial effects of the technical solutions provided in this application are:

[0026] This application provides a cloud-edge collaborative multimodal analysis method, applied to both cloud and edge terminals. The method applied to the cloud terminal includes: assigning analysis nodes of the analysis process corresponding to the analysis task to both the cloud terminal and the edge terminal; during the execution of the analysis process, after determining that the current analysis node is on the cloud terminal, if the current analysis node is not configured with a duplicate identifier, performing corresponding analysis operations, including analyzing multimodal data, to obtain the analysis results; then, after determining that the next analysis node is on the edge terminal, sending the analysis results to the edge terminal; otherwise, continuing the analysis operation of the next analysis node on the cloud terminal. This method achieves the goal of cloud-edge collaborative multimodal data analysis by comprehensively utilizing the resources of both cloud and edge terminals to process multimodal data analysis tasks. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0028] Figure 1 A schematic diagram of the architecture of a cloud-edge collaborative multimodal analysis system provided in this application embodiment;

[0029] Figure 2a A flowchart illustrating a merging process provided in an embodiment of this application;

[0030] Figure 2b A flowchart illustrating a resource scheduling process provided in an embodiment of this application;

[0031] Figure 3 A flowchart illustrating a cloud-edge collaborative multimodal analysis method provided in this application embodiment;

[0032] Figure 4 A flowchart illustrating another cloud-edge collaborative multimodal analysis method provided in this application embodiment;

[0033] Figure 5 A schematic diagram of the structure of a cloud-edge collaborative multimodal analysis device provided in this application embodiment;

[0034] Figure 6 A schematic diagram of another cloud-edge collaborative multimodal analysis device provided in this application embodiment. Detailed Implementation

[0035] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0036] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0038] First, let's introduce and explain several terms used in this application:

[0039] Multimodal biometrics refers to the integration or fusion of two or more biometric technologies. Leveraging the unique advantages of these technologies and combining them with data fusion techniques, it makes the authentication and identification process more accurate and secure. The main difference between multimodal biometrics and traditional single-biometric methods is that multimodal biometrics can use a single sensor (either independent or a combination of multiple acquisition methods) to collect different biometric features (such as fingerprints, veins, faces, iris images, etc.), and then analyze these features to determine the characteristic values ​​of each biometric method for identification and authentication.

[0040] With the continuous iteration of AI technology and information collection methods, the sources and processing methods of information are becoming increasingly diversified. Single recognition technologies are no longer sufficient to meet this development, and multimodal biometric technology is receiving increasing attention. A popular research direction in the development of multimodal biometric technology is the multimodal learning between images, videos, audio sources, and semantics. Currently, multimodal biometric technology is mainly applied in fields such as online entertainment, identity authentication, healthcare, smart finance, security, education, military industry, and park management.

[0041] Following the background technology, which describes three artificial intelligence solutions, specifically:

[0042] Single-modal approach: Currently, single-modal analysis models are used for intelligent analysis in most scenarios. However, this can only solve problems in simple, single-scenario situations. Single-modal models have low accuracy or cannot recognize complex and personalized scenarios.

[0043] Multimodal approach: Multimodal analysis model is an artificial intelligence technology that has only been studied in recent years. This approach is mainly used for multimodal recognition of speech and vision, but there are not many mature models.

[0044] Orchestration-based solutions: Multimodal analysis is achieved through process-based orchestration. However, its orchestration and analysis capabilities are relatively simple. It suffers from wasted computing power due to repeated execution, performance issues caused by not considering resource scheduling, and it does not support cloud-edge scenarios.

[0045] In comparison, orchestration-based solutions are significantly better than single-modal and multi-modal solutions, but the advantages are not obvious, and there are problems such as wasted computing power, lack of consideration for resource scheduling, and uncertainty about whether they can be applied to cloud-edge scenarios.

[0046] This application provides a multimodal analysis method, apparatus, electronic device, computer-readable storage medium, and computer program product based on cloud-edge collaboration, which aims to solve the above-mentioned technical problems in the prior art.

[0047] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0048] See Figure 1This application provides an architectural diagram of a cloud-edge collaborative multimodal analysis system. The system 100 includes a cloud terminal 110 and an edge terminal 120. The cloud terminal 110 mainly includes a process library 111, a process management engine 112, multiple executors 113 (specifically including executors 113a, 113b, etc.), and a resource scheduler 114. The executors 113 are also connected to a knowledge base 115 and an AI model processor 116. The process management engine 112 provides process parsing and process control functions. The edge terminal 120 includes a modal data processor 121 and multiple executors 122 (specifically including executors 122a, 122b, etc.). The executors 122 are also connected to a knowledge base 123 and an AI model processor 124. The process management engine 112 is also connected to the executors 122 to facilitate task assignment to the executors 122.

[0049] For example, the functions of each module and sub-module of system 100 will be illustrated below through the process of system 100 processing analysis task (1). Among them, the process of processing analysis task (1) mainly involves three processes: intelligent analysis process, resource scheduling process, and merging process when adding analysis task (1). Among them, analysis task (1) corresponds to analysis process (1).

[0050] See Figure 2a This application provides a flowchart of a merging process to illustrate the above-mentioned merging process in detail. When the cloud terminal 110 receives a new instruction carrying an analysis task (1), the process management engine 112 reads the analysis process (1) according to the analysis task (1) and performs a merging operation on all analysis nodes of the analysis process (1) starting from the starting analysis node of the analysis process (1).

[0051] Specifically, the characteristics of the current analysis node are obtained. For example, the characteristics of the current analysis node may include input parameters, processing procedures, and output results. Based on the analysis process set, it is determined whether there are analysis nodes with the same characteristics in the analysis process set. Specifically, the same characteristics refer to analysis nodes that have the same input parameters, the same processing procedures, and the same output results as the current analysis node. Such analysis nodes in the analysis process set are called matching analysis nodes. For the two scenarios of the existence or non-existence of matching analysis nodes, this application embodiment corresponds to a merging method. The analysis process set is formed by merging the analysis nodes of all historical analysis processes in the process library 111 according to a preset method.

[0052] If a matching analysis node exists, configure the same duplicate identifier for both the current analysis node and the matching analysis node in the analysis workflow set. If the matching analysis node has already been configured with a duplicate identifier, simply configure the same duplicate identifier for the current analysis node.

[0053] If no matching analysis node exists, the current analysis node is added to the analysis process set according to a preset method. This preset method may be: determining the routing information of the current analysis node, which includes the previous analysis node and the routing conditions from the previous analysis node to the current analysis node, where the previous analysis node is already in the analysis process set. Based on this routing information, the current analysis node is merged into the analysis process set and connected to the previous analysis node in the routing information.

[0054] After merging the current analysis node, the next analysis node in analysis process (1) can be merged into the analysis process set in the same way. For the sake of simplicity, it will not be described in detail here.

[0055] If the next analysis node is the termination analysis node, the merging process ends.

[0056] In addition, after receiving a new instruction carrying an analysis task (1), the analysis process (1) can be stored in the process library 111 and a corresponding identifier can be assigned to the analysis process (1).

[0057] See you again Figure 1 This application embodiment also provides an intelligent analysis process. Specifically, after receiving a start command carrying an analysis task (1), the cloud terminal 110 starts the analysis task (1).

[0058] First, the process management engine 112 reads the analysis process (1) from the process library 111. The process management engine 112 parses the analysis process (1) to determine the resource allocation scheme for each analysis node in the analysis process (1). Then, the process management engine 112 divides all the analysis nodes into two parts according to the resource allocation scheme for each analysis node: the part processed by the cloud terminal 110 and the part processed by the edge terminal. The process management engine 112 then allocates the analysis nodes to the corresponding terminals according to the resource allocation scheme.

[0059] The process management engine 112 starts the executors of the cloud terminal 110 and the edge terminal 120 to begin executing the analysis task (1).

[0060] After receiving the multimodal data, the modal data processor 121 performs preprocessing operations on the multimodal data to obtain the multimodal data to be analyzed. If the current analysis node of the analysis process (1) is located at the edge terminal 120, the modal data processor 121 sends the multimodal data to be analyzed to the executor 122. If the current analysis node of the analysis process (1) is located at the cloud terminal 110, the modal data processor 121 sends the multimodal data to be analyzed to the executor 113 of the cloud terminal 110. The multimodal data includes, but is not limited to, video, audio, text, and sensor-collected data. The preprocessing includes, but is not limited to, spatiotemporal alignment, for example, summarizing videos and audios containing the pet cat B to obtain a set, which can be the multimodal data to be analyzed; and collecting video, audio, text, and sensor data related to the entrance of the community into a set on a daily basis, which can be the multimodal data to be analyzed.

[0061] After receiving the multimodal data to be analyzed, executor 113 calls AI model processor 116 to analyze the multimodal data and obtain the analysis results; after receiving the multimodal data to be analyzed, executor 112 calls AI model processor 124 to analyze the multimodal data. Before executing the analysis operation corresponding to the analysis node, executors 113 and 122 can first determine whether the analysis node is configured with a duplicate identifier. If it is determined that a duplicate identifier is configured, they query the database of system 100 to see if there are any historical analysis results for that analysis node (i.e., analysis results of analysis nodes configured with the same duplicate identifier). If historical analysis results exist, the analysis operation is abandoned, and the historical analysis results are used as the analysis results.

[0062] After processing, once any executor obtains the analysis results, the process management engine 112 is also responsible for determining the next analysis node and sending the analysis results and / or the multimodal data to be analyzed to the next analysis node.

[0063] When the analysis process (1) reaches the last analysis node and the analysis result of the last analysis node is obtained, the cloud terminal 110 outputs the analysis result.

[0064] For example, after the property management of the residential community adopts System 100, the analysis tasks that can be input are: (A) security monitoring at the community entrance; (B) tracking the trajectory of pet cat B. For analysis task (A), video data collected by cameras, sensors, etc., near the community can be collected to obtain multimodal data to be analyzed. The multimodal data is then analyzed and processed according to the analysis process corresponding to analysis task (A) to obtain specific information about the security monitoring at the community entrance, such as pedestrian flow information and information on outsiders entering the community (delivery drivers, couriers, etc.). For analysis task (B), video and audio data containing pet cat B can be collected to obtain multimodal data to be analyzed. The multimodal data is then analyzed according to the analysis process corresponding to analysis task (B) to obtain trajectory information about pet cat B. If pet cat B gets lost in the community, analysis task (B) can also provide valuable information to help find pet cat B.

[0065] See Figure 2b This application embodiment also provides a flowchart of the resource scheduling process. After the analysis task (1) is started, the process management engine 112 reads the analysis process (1) from the process library 111 according to the analysis task (1). The process management engine 112 parses the analysis process (1) to determine the resource allocation scheme of each analysis node in the analysis process (1).

[0066] Specifically, firstly, the resource consumption of analysis process (1) is estimated by estimating the resource amount required by each analysis node in analysis process (1). After estimating the resource amount required by each analysis node, the resource consumption of analysis process (1) is calculated. The current remaining resources of system 100 (the sum of the remaining resources of all executors of cloud terminal 110 and edge terminal 120) are obtained; the resource consumption of analysis process (1) and the size of the current remaining resources are determined. If the resource consumption is not greater than the current remaining resources, a matching executor (which can be executor 113 or executor 122) is further found for each analysis node in analysis process (1). The matching criterion is that the remaining resources of the executor are not less than the resource amount required by the analysis node.

[0067] In addition, a unified unit can be used to specifically describe the amount of resources. One executor of cloud terminal 110 can provide 100 computing power resources, of which 60 computing power is currently occupied, leaving 40 computing power remaining. Here, the executor can be a server, and different models and configurations of servers correspond to different computing power. Cloud terminal 110 and edge terminal 120 can be a server cluster.

[0068] Continue to determine whether all analysis nodes in the analysis process (1) have a matching executor. If each analysis node has a matching executor, then determine each matching executor and the terminal to which the matching executor belongs as the allocation scheme for the corresponding analysis node, and add the analysis task (1) to the task scheduling queue. If only some analysis nodes have matching executors, then the matching is not identified, and the analysis task (1) cannot be added to the task scheduling queue.

[0069] See Figure 3 This application also provides a multimodal analysis method based on cloud-edge collaboration, which can be applied to the cloud terminal 110 in the system 100 of the above embodiment. The method includes:

[0070] S310 assigns the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal.

[0071] An analysis process includes at least one analysis node. The analysis task can be the analysis task (1) in the above embodiments, and the analysis process can be the analysis process (1).

[0072] S320: After determining that the current analysis node in the analysis process is located on the cloud terminal, if the current analysis node is not configured with a duplicate identifier, the analysis operation corresponding to the current analysis node is executed to obtain the analysis results. The analysis operation includes analyzing the multimodal data to be analyzed through an AI model.

[0073] S330: If the next analysis node of the current analysis node is at the edge terminal, send the analysis results to the edge terminal.

[0074] This application provides a cloud-edge collaborative multimodal analysis method, applied to both cloud and edge terminals. The method applied to the cloud terminal includes: assigning analysis nodes of the analysis process corresponding to the analysis task to both the cloud terminal and the edge terminal; during the execution of the analysis process, after determining that the current analysis node is on the cloud terminal, if the current analysis node is not configured with a duplicate identifier, performing corresponding analysis operations, including analyzing multimodal data, to obtain the analysis results; then, after determining that the next analysis node is on the edge terminal, sending the analysis results to the edge terminal; otherwise, continuing the analysis operation of the next analysis node on the cloud terminal. This method comprehensively utilizes the resources of both cloud and edge terminals to process multimodal data analysis tasks, achieving the goal of cloud-edge collaborative processing of multimodal data analysis tasks.

[0075] This application embodiment also provides a possible implementation method, which further includes, before allocating the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal respectively:

[0076] Read the analysis process from the database; parse the analysis process to determine the allocation scheme for each analysis node in the analysis process;

[0077] Specifically, the analysis nodes of the analysis process corresponding to the analysis task are allocated to cloud terminals and edge terminals respectively, including: allocating the analysis nodes of the analysis process to cloud terminals or edge terminals according to the allocation scheme of each analysis node.

[0078] In one possible implementation, both the cloud terminal and the edge terminal include at least one actuator that parses and analyzes the process to determine the allocation scheme for each analysis node in the analysis process, including:

[0079] Calculate the resource consumption of the analysis process; if the resource consumption is not greater than the current remaining resources, execute the matching operation corresponding to each analysis node in the analysis process to determine the executor that matches each analysis node.

[0080] The current remaining resources are the sum of the remaining resources of all executors in the cloud terminal and edge terminal; if all matching operations result in a successful match, each matching executor and its associated terminal will be determined as the allocation scheme for the corresponding analysis node.

[0081] The resource consumption of this analysis process is the sum of the estimated resources required for each analysis node. These resources include, but are not limited to, CPU resources, GPU resources, and memory resources.

[0082] The executor is executor 113 in the above embodiment. In practical applications, the executor can be a server, and the cloud terminal and edge terminal can be a server cluster.

[0083] This application embodiment also provides a possible implementation method, which further includes, before allocating the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal:

[0084] If a new instruction carrying an analysis process is received, perform the following operations for each analysis node in the analysis process: Perform a matching operation on the analysis nodes according to the historical analysis process set, where the historical analysis process set is formed by merging the analysis nodes of all historical analysis processes according to a preset method.

[0085] If a match is successful, a duplicate identifier is configured for the analysis node and the matching analysis node, wherein the analysis node and the matching analysis node have the same input, the same processing procedure and the same output; if a match fails, the analysis node is merged into the historical analysis process set according to a preset method.

[0086] The preset method can be referred to in the above embodiments, and will not be repeated here for the sake of simplicity.

[0087] In one possible implementation, after determining that the current analysis node in the analysis process is located on a cloud terminal, the method further includes:

[0088] If the current analysis node has a duplicate identifier configured, query the historical analysis results of other analysis nodes that have the same duplicate identifier configured as the analysis node; if the historical analysis results are not empty, abandon the analysis operation corresponding to the current analysis node, and use the historical analysis results as the analysis results of the current analysis node.

[0089] Specifically, a query command carrying a duplicate identifier is sent to the process management engine in the cloud terminal. The process management engine queries the database for historical analysis results based on the duplicate identifier. If the duplicate results exist, the historical analysis results are sent to the corresponding executor.

[0090] In one possible implementation, the method also includes storing the analysis results, specifically:

[0091] If the analysis result was obtained after performing the analysis operation corresponding to the current analysis node, check whether the current analysis node has a duplicate flag configured. If the check result indicates that a duplicate flag has been configured, store the analysis result according to the duplicate flag.

[0092] or,

[0093] If a storage request carrying analysis results and corresponding duplicate identifiers is received from an edge terminal, the analysis results are stored according to the corresponding duplicate identifiers.

[0094] See Figure 4 This application also provides a multimodal analysis method based on cloud-edge collaboration, which can be applied to the edge terminal 120 in the above embodiments. The method includes:

[0095] S410 receives the analysis node assigned by the cloud terminal, which corresponds to the analysis task and the analysis process.

[0096] S420: After determining that the current analysis node in the analysis process is at the edge terminal, if the current analysis node is not configured with a duplicate identifier, execute the analysis operation corresponding to the current analysis node to obtain the analysis result. The analysis operation includes at least analyzing the multimodal data to be analyzed through an AI model.

[0097] The edge terminal includes multiple actuators, which perform the analysis operations corresponding to the current analysis node in the domain. This actuator can be actuator 122 as described in the above embodiment.

[0098] In one possible implementation, the method may further include:

[0099] If the current analysis node has a duplicate identifier configured, query the historical analysis results of other analysis nodes that have the same duplicate identifier configured as the analysis node; if the historical analysis results are not empty, abandon the analysis operation corresponding to the current analysis node, and use the historical analysis results as the analysis results of the current analysis node.

[0100] Specifically, a query message carrying a duplicate identifier is sent to the cloud terminal, and the cloud terminal retrieves historical analysis results from the database based on the duplicate identifier.

[0101] S430: If the next analysis node of the current analysis node is in the cloud terminal, send the analysis results to the cloud terminal.

[0102] In one possible implementation, the method further includes:

[0103] If the analysis result is obtained after performing the analysis operation corresponding to the current analysis node, check whether the current analysis node is configured with a duplicate flag.

[0104] If the check result indicates that a duplicate identifier has been configured, a storage request carrying the analysis results and the corresponding duplicate identifier is sent to the cloud terminal.

[0105] This application embodiment also provides a possible implementation, in which, after receiving at least one analysis node of the analysis process corresponding to the analysis task allocated by the cloud terminal, the method further includes:

[0106] Receive raw multimodal data; preprocess the raw multimodal data to obtain the multimodal data to be analyzed; if the initial node of the analysis process is determined to be in the cloud terminal, feed back the multimodal data to be analyzed to the cloud terminal.

[0107] The preprocessing procedure and multimodal data can be referred to in the above embodiments, and will not be repeated here for the sake of simplicity.

[0108] See Figure 5 This application also provides a cloud-edge collaborative multimodal analysis device, which can be applied to the cloud terminal 110 in the above embodiments. The device 500 includes:

[0109] The allocation module 510 is used to allocate the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal.

[0110] The execution module 520 is used to perform the analysis operation corresponding to the current analysis node after determining that the current analysis node of the analysis process is in the cloud terminal, if the current analysis node is not configured with a duplicate identifier, and obtain the analysis result. The analysis operation includes analyzing the multimodal data to be analyzed through an AI model.

[0111] The transceiver module 530 is used to send the analysis results to the edge terminal if the next analysis node of the current analysis node is located at the edge terminal.

[0112] In one possible implementation, the device 500 includes a parsing module 540, which, before assigning the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal respectively, is specifically used for:

[0113] Read the analysis process from the database; parse the analysis process to determine the allocation scheme for each analysis node in the analysis process;

[0114] Specifically, the analysis nodes of the analysis process corresponding to the analysis task are allocated to cloud terminals and edge terminals respectively, including: allocating the analysis nodes of the analysis process to cloud terminals or edge terminals according to the allocation scheme of each analysis node.

[0115] In one possible implementation, the cloud terminal and the edge terminal each include at least one actuator, and the parsing module 540, in parsing the analysis process to determine the allocation scheme for each analysis node in the analysis process, is specifically used for:

[0116] Calculate the resource consumption of the analysis process;

[0117] If the resource consumption is not greater than the current remaining resources, perform the matching operation corresponding to each analysis node in the analysis process to determine the executor that matches each analysis node. The current remaining resources are the sum of the remaining resources of all executors in the cloud terminal and the edge terminal.

[0118] If all matching operations result in a successful match, the matching executor and its associated terminal are determined as the allocation scheme for the corresponding analysis node.

[0119] In one possible implementation, the device 500 further includes a merging module 550, specifically used for:

[0120] If a new instruction carrying an analysis flow is received, perform the following operations for each analysis node in the analysis flow:

[0121] Matching operations are performed on the analysis nodes based on the historical analysis process set, which is formed by merging the analysis nodes of all historical analysis processes according to a preset method.

[0122] If a match is successful, a duplicate identifier is configured for the analysis node and the matching analysis node, where the analysis node and the matching analysis node have the same input, the same processing procedure and the same output;

[0123] If a match fails, the analysis node will be merged into the historical analysis process set according to the preset method.

[0124] In one possible implementation, after determining that the current analysis node of the analysis process is located on the cloud terminal, the execution module 520 can also be used for:

[0125] If the current analysis node has a duplicate identifier configured, query the historical analysis results of other analysis nodes that have the same duplicate identifier configured as the current analysis node;

[0126] If the historical analysis result is not empty, abandon the analysis operation corresponding to the current analysis node, and use the historical analysis result as the analysis result of the current analysis node.

[0127] See Figure 6 This application also provides a cloud-edge collaborative multimodal analysis device for use in edge terminals. The device 600 includes:

[0128] The transceiver module 610 is used to receive the analysis nodes of the analysis process corresponding to the analysis task allocated by the cloud terminal;

[0129] The execution module 620 is used to perform the analysis operation corresponding to the current analysis node after determining that the current analysis node in the analysis process is at the edge terminal, if the current analysis node is not configured with a duplicate identifier, and obtain the analysis result. The analysis operation includes at least analyzing the multimodal data to be analyzed through an AI model.

[0130] The transceiver module 610 is used to send the analysis results to the cloud terminal if the next analysis node of the current analysis node is located in the cloud terminal.

[0131] In one possible implementation, the transceiver module 610 is also used for:

[0132] Receive raw multimodal data; preprocess the raw multimodal data to obtain the multimodal data to be analyzed; if the initial node of the analysis process is determined to be in the cloud terminal, feed back the multimodal data to be analyzed to the cloud terminal.

[0133] This application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of any of the cloud-edge collaborative multimodal analysis methods shown in the above embodiments of this application.

[0134] Electronic devices include, but are not limited to, servers.

[0135] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the cloud-edge collaborative multimodal analysis methods shown in the above embodiments of this application.

[0136] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps of any of the cloud-edge collaborative multimodal analysis methods shown in the above embodiments of this application.

[0137] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.

[0138] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0139] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.

Claims

1. A multimodal analysis method based on cloud-edge collaboration, characterized in that, Applied to cloud terminals, the method includes: Assign analysis nodes of the analysis process corresponding to the analysis task to cloud terminals and edge terminals; After determining that the current analysis node of the analysis process is located on the cloud terminal, if the current analysis node is not configured with a duplicate identifier, the analysis operation corresponding to the current analysis node is executed to obtain the analysis result; if the current analysis node has configured the duplicate identifier, the historical analysis results of other analysis nodes configured with the same duplicate identifier are queried; if the historical analysis results are not empty, the analysis operation corresponding to the current analysis node is abandoned, and the historical analysis results are used as the analysis result of the current analysis node; wherein, the analysis operation includes analyzing the multimodal data to be analyzed through an AI model; the multimodal data includes video, audio, text, and data collected by sensors; the multimodal data to be analyzed is obtained by preprocessing the multimodal data, and the preprocessing includes spatiotemporal alignment; If the next analysis node of the current analysis node is located at the edge terminal, the analysis result is sent to the edge terminal.

2. The method according to claim 1, characterized in that, Before assigning the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal respectively, the method further includes: The analysis process is read from the database; The analysis process is analyzed to determine the allocation scheme for each analysis node in the analysis process; The method of allocating the analysis nodes of the analysis process corresponding to the analysis task to the cloud terminal and the edge terminal respectively includes: According to the allocation scheme of each analysis node, the analysis nodes of the analysis process are allocated to cloud terminals or edge terminals.

3. The method according to claim 2, characterized in that, The cloud terminal and the edge terminal each include at least one executor. The step of parsing the analysis process to determine the allocation scheme for each analysis node in the analysis process includes: Calculate the resource consumption of the analysis process; If the resource consumption is not greater than the current remaining resources, a matching operation corresponding to each analysis node in the analysis process is executed to determine the executor that matches each analysis node, wherein the current remaining resources are the sum of the remaining resources of all executors of the cloud terminal and the edge terminal; If all matching operations result in a successful match, the matching executor and its associated terminal are determined as the allocation scheme for the corresponding analysis node.

4. The method according to any one of claims 2-3, characterized in that, Before reading the analysis process from the database, the method further includes: If a new instruction carrying the analysis process is received, perform the following operations for each analysis node in the analysis process: A matching operation is performed on the analysis nodes according to the historical analysis process set, wherein the historical analysis process set is formed by merging the analysis nodes of all historical analysis processes in a preset manner; If a match is successful, the duplicate identifier is configured for the analysis node and the matching analysis node, wherein the analysis node and the matching analysis node have the same input, the same processing procedure and the same output; If a match fails, the analysis node will be merged into the historical analysis process set according to the preset method.

5. A multimodal analysis method based on cloud-edge collaboration, characterized in that, Applied to edge terminals, the method includes: Receive the analysis nodes of the analysis process corresponding to the analysis task assigned by the cloud terminal; After determining that the current analysis node in the analysis process is located at the edge terminal, if the current analysis node is not configured with a duplicate identifier, the analysis operation corresponding to the current analysis node is executed to obtain the analysis result; if the current analysis node has configured the duplicate identifier, the historical analysis results of other analysis nodes configured with the same duplicate identifier are queried; if the historical analysis results are not empty, the analysis operation corresponding to the current analysis node is abandoned, and the historical analysis results are used as the analysis result of the current analysis node; wherein, the analysis operation includes at least analyzing the multimodal data to be analyzed through an AI model; the multimodal data includes video, audio, text, and data collected by sensors; If the next analysis node of the current analysis node is in the cloud terminal, the analysis result is sent to the cloud terminal.

6. The method according to claim 5, characterized in that, After receiving at least one analysis node of the analysis process corresponding to the analysis task allocated by the cloud terminal, the method further includes: Receive raw multimodal data; The original multimodal data is preprocessed to obtain the multimodal data to be analyzed; If the initial node of the analysis process is determined to be in the cloud terminal, the multimodal data to be analyzed is fed back to the cloud terminal.

7. A multimodal analysis device based on cloud-edge collaboration, characterized in that, The device, applied to a cloud terminal, includes: The allocation module is used to allocate the analysis nodes of the analysis process corresponding to the analysis task to cloud terminals and edge terminals; The execution module is configured to, after determining that the current analysis node of the analysis process is located on the cloud terminal, if the current analysis node is not configured with a duplicate identifier, execute the analysis operation corresponding to the current analysis node to obtain the analysis result; if the current analysis node is configured with the duplicate identifier, query the historical analysis results of other analysis nodes configured with the same duplicate identifier; if the historical analysis results are not empty, abandon the execution of the analysis operation corresponding to the current analysis node, and use the historical analysis results as the analysis result of the current analysis node; wherein, the analysis operation includes analyzing the multimodal data to be analyzed through an AI model; the multimodal data includes video, audio, text, and sensor-collected data; the multimodal data to be analyzed is obtained by preprocessing the multimodal data, the preprocessing including spatiotemporal alignment; The transceiver module is used to send the analysis results to the edge terminal if the next analysis node of the current analysis node is located at the edge terminal.

8. A multimodal analysis device based on cloud-edge collaboration, characterized in that, The device, applied to an edge terminal, includes: The transceiver module is used to receive analysis nodes of the analysis process corresponding to the analysis task allocated by the cloud terminal; An execution module is configured to, after determining that the current analysis node of the analysis process is located at the edge terminal, if the current analysis node is not configured with a duplicate identifier, execute the analysis operation corresponding to the current analysis node to obtain the analysis result; if the current analysis node is configured with the duplicate identifier, query the historical analysis results of other analysis nodes configured with the same duplicate identifier; if the historical analysis results are not empty, abandon the execution of the analysis operation corresponding to the current analysis node, and use the historical analysis results as the analysis result of the current analysis node; wherein, the analysis operation includes at least analyzing the multimodal data to be analyzed using an AI model; the multimodal data includes video, audio, text, and data collected by sensors; the multimodal data to be analyzed is obtained by preprocessing the multimodal data, the preprocessing including spatiotemporal alignment; The transceiver module is used to send the analysis results to the cloud terminal if the next analysis node of the current analysis node is located in the cloud terminal.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-4 or 5-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-4 or 5-6.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-4 or 5-6.

Citation Information

Patent Citations

  • Cross-layer sensory self-configuration system and method based on multi-scale entropy

    CN103440224A

  • Space-time data visualization task execution method based on cloud edge-end architecture

    CN113032132A

  • Cloud side-end collaborative resource management method and system

    CN113452566A