AI multi-mode fusion interaction method, device, system and equipment

By identifying user voice information and switching scene modes through edge devices and triggering the AI ​​function module to process multimodal information, the problems of low interaction efficiency and delay in information processing under the requirements of multiple scenarios in the prior art are solved, and efficient multimodal information processing and interactive intelligence are achieved.

CN120179079AActive Publication Date: 2025-06-20XIAMEN RGBLINK SCI & TECH CO LTD

Patent Information

Application Number
CN202510647041.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing multimodal interaction technology is difficult to flexibly switch scenario modes under the needs of multiple scenarios, and the multimodal information processing lacks an effective collaborative working mechanism, resulting in low interaction efficiency, delay in information processing and inaccurate output.

Method used

Obtain user voice information through edge devices, identify key information and switch scene modes, and trigger the corresponding AI function module to process multimodal information. If the edge device cannot meet the processing requirements, the result data is sent to the cloud service device for priority processing, generate multimodal fusion data and return to the edge device.

Benefits of technology

It realizes the flexibility to switch scene modes under the needs of multiple scenarios to improve interaction efficiency; through the collaboration between edge devices and cloud service devices, efficient processing and integration of multimodal information, optimize resource utilization, and realize intelligent and personalized interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179079A_ABST
    Figure CN120179079A_ABST
Patent Text Reader

Abstract

The invention discloses an AI multi-mode fusion interaction method, device, system and equipment. An edge device obtains multi-mode information of a user in a current scene mode; preprocessing the multi-modal information, and outputting result data conforming to the current scene mode; when the edge device opens the uploading authority and the AI function module cannot meet the multi-modal information processing requirement, the result data is sent to the cloud service device, so that the cloud service device carries out processing according to the priority processing strategy of the multi-modal information to generate multi-modal fusion data, and the multi-modal fusion data is returned to the edge device for storage. And the edge device determines whether to publish the multi-modal fusion data according to the service requirement and synchronizes the multi-modal fusion data to the mobile terminal device. According to the method and the device, the corresponding scene mode can be flexibly switched according to different scene requirements, and each AI function module is integrated on the edge equipment, so that the corresponding AI function module is triggered to process the multi-modal information and send the multi-modal information to the cloud service equipment for fusion processing, and efficient processing and fusion of the multi-modal information are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing, and in particular to an AI multimodal fusion interaction method, device, system and equipment. Background Art

[0002] In the existing multimodal interaction technologies, although there are already some technologies related to speech recognition, in the context of the rapid development of information technology today, multimodal interaction technologies have been widely used in many fields. However, there are still many problems to be solved in the existing technologies. On the one hand, most of the current multimodal interaction methods can only be optimized for a single specific scenario. Under the requirements of multiple scenarios, it is difficult to flexibly switch to the corresponding scenario mode according to different scenario requirements, resulting in users having to perform cumbersome operations to adjust the system state, which greatly reduces the interaction efficiency.

[0003] On the other hand, there are also obvious deficiencies in the processing of multimodal information. When the existing technologies process various types of information such as sound, audio and images, the various AI function modules in the system are often independent of each other and lack an effective collaborative working mechanism. This leads to the difficulty of calling the corresponding AI function modules to perform fusion processing on multimodal information based on the scenario requirements when facing complex user requirements. It may not only lead to misunderstandings of user intentions or omissions of information, but also affect the accuracy and real-time performance of the final output results, thus affecting the interaction effect. For example, in a task involving both voice communication and image recognition, it may not be possible to determine the processing priority of sound information and image information, resulting in the delay of important information processing and affecting the user's timely acquisition and use of the interaction results. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides the following technical solutions: An AI multimodal fusion interaction method, applied to an edge device, is characterized by comprising: Obtaining the first voice information of a user; Identifying the key information of the first voice information and switching to the corresponding scenario mode based on the key information; Triggering at least one AI function module based on the current scenario mode to obtain the multimodal information of the user in the current scenario mode; Preprocessing the multimodal information and outputting result data that conforms to the current scenario mode according to a preset rule; When the edge device enables the upload permission and the AI function module cannot meet the requirements of the multimodal information processing, the result data is sent to the cloud service device, so that the cloud service device processes it according to the priority processing strategy of the multimodal information to generate multimodal fusion data, returns it to the edge device for storage, and decides whether to publish the multimodal fusion data according to the service requirements; Synchronize the multimodal fusion data and its publication information to the mobile terminal device; Among them, the AI function module includes at least one of an AI auditory module, an AI speech module, an AI vision module, and an AI creation module; the multimodal information includes sound information, audio information, and image information; among them, the priority processing strategy of the multimodal information is determined based on the current scene mode.

[0005] As a further improvement, the edge device establishes a connection with the cloud service device through an adaptive communication protocol. When the edge device is disconnected from the cloud service device, the edge device continues to process the multimodal information through the locally pre-stored AI function module, and automatically synchronizes the data to the cloud service device when reconnecting.

[0006] As a further improvement, when the edge device enables the upload permission and the AI function module cannot meet the requirements of the multimodal information processing, sending the result data to the cloud service device includes: after the edge device establishes a connection with the mobile terminal device, the mobile terminal device will pre-configure the upload permission for the edge device to send the result data to the cloud service device, so that when the edge device enables the upload permission, the result data is sent to the cloud service device.

[0007] As a further improvement, when the scene mode is a meeting scene mode, the AI auditory module, the AI speech module, the AI vision module, and the AI creation module are triggered based on the current meeting scene mode; The AI speech module synchronously outputs translated audio according to the requirements, and adjusts the speech emotion intensity according to the gesture amplitude and expression of the speaker captured by the AI vision module; The AI auditory module obtains the sound information of the speaker; The AI vision module scans the image information in the shared screen to generate text information; The AI creation module automatically summarizes the meeting minutes after the meeting ends.

[0008] An AI multimodal fusion interaction method, applied to a cloud service device, includes: Obtain the result data sent by at least one edge device; the result data is that after the edge device obtains the first voice information of the user, switches to the corresponding scenario mode according to the key information of the first voice information, triggers at least one AI function module in the current scenario mode to obtain the multimodal information of the user in the current scenario mode, preprocesses the multimodal information, and outputs the result data that conforms to the current scenario mode according to the preset rules; Among them, the AI function module includes at least one of an AI auditory module, an AI voice module, an AI visual module, and an AI creation module; the multimodal information includes sound information, audio information, and image information; Call the AI large model to perform fusion processing on the result data. The AI large model matches the priority processing strategy of the pre-stored multimodal information according to the scenario mode in the result data, and processes each modal information in the result data according to the priority processing strategy of the multimodal information, and outputs multimodal fusion data; among them, the priority processing strategy sets the priority according to the scenario mode and modal information; Return the multimodal fusion data to the edge device for storage, so that the edge device decides whether to publish the multimodal fusion data according to the service requirements, and synchronize the multimodal fusion data and its publishing information to the mobile terminal device.

[0009] As a further improvement, when processing each modal information in the result data according to the priority processing strategy of the multimodal information and outputting multimodal fusion data, it includes: Match the corresponding modal computing power allocation rule according to the priority processing strategy of each modal information, and process each modal information according to the modal computing power allocation rule, and output multimodal fusion data; among them, the modal computing power allocation rule is set based on the computing power requirements of each modal information in the current scenario mode to match the computing power requirements of each modal information processing.

[0010] As a further improvement, calling the AI large model to perform fusion processing on the result data includes: When obtaining multiple result data sent by multiple edge devices, determine the scenario mode of each result data, and process each result data according to the scenario computing power allocation rule; among them, the scenario computing power allocation rule is set based on the security requirements, real-time requirements, and general requirements of multiple scenario modes to match the computing power requirements of each scenario mode.

[0011] As a further improvement, when the scene mode is a meeting scene, the audio information, voice information, and image information in the multimodal information are processed in sequence according to the priority processing strategy, and computing power is allocated to the audio information, voice information, and image information in sequence according to the modal computing power allocation rule.

[0012] An AI multimodal fusion interaction edge device, comprising: A data recognition module, configured to obtain the first voice information of a user, recognize the key information of the first voice information, and switch to the corresponding scene mode based on the key information; An AI function call module, configured to trigger at least one AI function module based on the current scene mode, where the AI function module includes an AI auditory module, an AI voice module, an AI vision module, and an AI creation module; A preprocessing module, configured to obtain the multimodal information of the user in the current scene mode, where the multimodal information includes voice information, audio information, and image information; preprocess the multimodal information, and output result data that conforms to the current scene mode according to a preset rule; A communication module, configured to, when the edge device enables the upload permission and the AI function module cannot meet the multimodal information processing requirement, send the result data to a cloud service device, so that the cloud service device processes the result data according to the priority processing strategy of the multimodal information to generate multimodal fusion data, return the multimodal fusion data to the edge device for storage, and decide whether to publish the multimodal fusion data according to service requirements; synchronize the multimodal fusion data and its publishing information to a mobile terminal device; where the priority processing strategy of the multimodal information is determined based on the current scene mode.

[0013] An AI-based multimodal data processing interaction ecosystem, comprising: An edge device, configured to obtain the first voice information of a user and recognize the key information of the first voice information, and switch to the corresponding scene mode based on the key information; Based on the current scene mode, trigger at least one AI function module, where the AI function module includes at least one of an AI auditory module, an AI voice module, an AI vision module, and an AI creation module; Obtain the multimodal information of the user in the current scene mode, where the multimodal information includes voice information, audio information, and image information; preprocess the multimodal information, and output result data that conforms to the current scene mode, and when the edge device enables the upload permission and the AI function module cannot meet the multimodal information processing requirement, send the result data to a cloud service device; A cloud service device obtains result data sent by at least one of the edge devices; Call an AI large model to perform fusion processing on the result data. The AI large model matches the priority processing strategy of the pre-stored multimodal information according to the scenario pattern in the result data, processes each modal information in the result data according to the priority processing strategy of the multimodal information, outputs multimodal fusion data, and returns it to the edge device for storage, so that the edge device decides whether to publish the multimodal fusion data according to business requirements; wherein, the priority processing strategy sets priorities according to the scenario pattern and modal information; A mobile terminal device receives the multimodal fusion data and its publication information synchronized by the edge device.

[0014] A computing device includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute any of the methods described above.

[0015] The beneficial effects are as follows: The present invention provides an AI multimodal fusion interaction method. The edge device switches to the corresponding scenario mode by obtaining the user's first voice information and identifying key information, and then triggers the corresponding AI function module to process the multimodal information. The preprocessed multimodal information outputs result data according to a preset rule. When the edge device enables the upload permission and the AI function module cannot meet the requirements of multimodal information processing, the result data is sent to the cloud service device. The cloud server device generates multimodal fusion data according to the priority processing strategy of the multimodal information and returns it to the edge device for storage. The edge device determines whether to publish the multimodal fusion data according to the actual business requirements, and synchronizes the multimodal fusion data and its publication information to the mobile terminal device. Under the requirements of multiple scenarios, the edge device can flexibly switch to the corresponding scenario mode according to different scenario requirements, improving the interaction efficiency; moreover, the edge device integrates various AI function modules, and can trigger the corresponding AI function module to preprocess the multimodal information according to the scenario requirements and send it to the cloud service device for fusion processing according to the priority processing strategy, which not only realizes the efficient processing and fusion of multimodal information, but also optimizes the utilization of resources, realizing the intelligent and personalized level of interaction. In the present invention, the efficient cooperation of the edge device, the cloud service device and the mobile device jointly completes the fusion processing of multimodal information, improves the data processing efficiency, and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of an AI multimodal fusion interaction method provided by an embodiment of the present invention when applied to an edge device; Figure 2 Flowchart of an AI multi-modal fusion interaction method provided by an embodiment of the present invention when applied to a cloud service device; Figure 3 Schematic diagram of modules of an AI multi-modal fusion interaction device provided by an embodiment of the present invention; Figure 4 Schematic diagram of modules of an AI multi-modal fusion interaction system provided by an embodiment of the present invention. Detailed implementation manners

[0017] In order to make the technical means, creative features, achieved purposes and functions of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments. However, the following embodiments are only the preferred embodiments of the present invention, not all of them. Based on the embodiments in the implementation manners, other embodiments obtained by those skilled in the art without creative efforts all belong to the protection scope of the present invention.

[0018] The present invention provides an AI multi-modal fusion interaction method. The edge device switches to the corresponding scene mode by obtaining the user's first voice information and identifying key information, and then triggers the corresponding AI function module to process multi-modal information. The pre-processed multi-modal information outputs result data according to preset rules, and then sends the result data to the cloud service device. The cloud server device generates multi-modal fusion data according to the priority processing strategy of the multi-modal information and returns it to the edge device for storage. The edge device determines whether to publish the multi-modal fusion data according to the actual business requirements, and synchronizes the multi-modal fusion data and its publishing information to the mobile terminal device. Under the requirements of multiple scenarios, the edge device can flexibly switch to the corresponding scene mode according to different scene requirements, improving the interaction efficiency. Moreover, the edge device integrates various AI function modules, and can trigger the corresponding AI function module to perform pre-processing and then send it to the cloud service device for fusion processing of multi-modal information according to the priority processing strategy, which not only realizes the efficient processing and fusion of multi-modal information, but also optimizes the utilization of resources, and realizes the intelligent and personalized level of interaction. The efficient cooperation of the edge device, cloud service device and mobile device in the present invention jointly completes the fusion processing of multi-modal information, improves the data processing efficiency, and improves the user experience.

[0019] The solution of the present application is applicable to a variety of scene modes, including but not limited to domestic conference scenes, cross-border conference scenes, content creation scenes, smart home scenes, intelligent driving scenes, medical consultation scenes and online education scenes, and all are applicable to the AI multi-modal fusion interaction method mentioned in the present application. Embodiment

[0020] Please refer to Figure 1, an AI multimodal fusion interaction method, applied to edge devices, includes: S101, obtaining the first voice information of the user; S102, identifying the key information of the first voice information, and switching to the corresponding scene mode based on the key information, where the key information includes execution action information, scene information, or device information; S103, triggering at least one AI function module based on the current scene mode, and the AI function module includes at least one of an AI auditory module, an AI voice module, an AI visual module, and an AI creation module; S104, obtaining the multimodal information of the user in the current scene mode, where the multimodal information includes sound information, audio information, and image information; S105, preprocessing the multimodal information and outputting result data that conforms to the current scene mode according to a preset rule; S106, when the edge device enables the upload permission and the AI function module cannot meet the multimodal information processing requirements, sending the result data to the cloud service device, so that the cloud service device processes and generates multimodal fusion data according to the priority processing strategy of the multimodal information, returns it to the edge device for storage, and decides whether to publish the multimodal fusion data according to the service requirements; among them, the priority processing strategy of the multimodal information is determined based on the current scene mode.

[0021] Synchronize the multimodal fusion data and its release information to the mobile terminal device.

[0022] Through the collaborative architecture of multimodal preprocessing of edge devices and high-order AI fusion of cloud service devices, the dual optimization of data processing efficiency and security is achieved. The edge device performs real-time collection, scene mode recognition, and preprocessing of multimodal information such as voice and images based on local computing power, generates result data, significantly reduces the cloud transmission load, and improves the response speed; when the AI function module in the edge device cannot meet the multimodal information processing requirements, the cloud service device relies on the computing power of the AI large model to perform in-depth fusion analysis on the result data, and realizes intelligent optimization through dynamic correction and cross-modal generation, forming an end-to-end efficient processing link.

[0023] In addition, for the privatization requirements of enterprise users, the edge device can independently complete high-frequency scene tasks (such as real-time translation of international conferences and digital human sample collection), and only upload non-sensitive data or tasks that require large computing power to the cloud service device. This can not only effectively reduce the time-consuming of the overall workflow and improve efficiency, but also collect user key data locally. Sensitive data can be preferentially stored in the local hard disk of the edge device and only uploaded to the cloud service device after authorization, effectively avoiding the risk of core data leakage and effectively protecting user privacy and data sovereignty.

[0024] In an embodiment of the present application, an AI multimodal fusion interaction method is applied to an edge device. The method is mainly described by taking a meeting scenario as an example, and is specifically as follows: Step S101: Obtain the first voice information of the user; The edge device is configured with a high-sensitivity microphone array to obtain the user's voice in real time. The microphone picks up sounds omnidirectionally and without dead angles, captures the user's voice completely, and performs digital processing on the captured sound, converting the analog signal into a digital signal. By skillfully applying filtering technology, the noisy sounds in the environment are removed, making the user's voice commands clearer and purer.

[0025] Step S102: Identify the key information of the first voice information, and switch to the corresponding scenario mode based on the key information. The key information includes execution action information, scenario information, or device information; The edge device pre-stores multiple scenario modes and their corresponding key information features. The scenario modes can be domestic meeting scenarios, international meeting scenarios, content creation scenarios, smart home scenarios, intelligent driving scenarios, medical consultation scenarios, or online education scenarios. The switching logic module of the scenario compares the extracted key information with the preset features. When they match, the corresponding scenario mode is triggered. Among them, the key information refers to the real-time information extracted from the user's voice commands, and the preset features refer to the feature information predefined and stored in the edge device, which correspond to each scenario mode and are used to identify and distinguish different scenario modes. The edge device determines whether to trigger the corresponding scenario mode by comparing the extracted key information with the preset features, so as to achieve flexible and intelligent scenario switching and meet the user's needs. Among them, the execution action information includes turn on, turn off, do, etc. The scenario information includes meeting, content creation, smart home, or intelligent driving, etc. The device information includes the device identification information in each scenario, such as the camera device identification and microphone identification in the meeting scenario. For example, in the meeting scenario, the voice information is "I want to start a meeting". The key information such as "start" and "meeting" is recognized, and the scenario is switched to the meeting scenario mode to prepare for recording and storing the meeting content, ensuring the efficient operation of the voice recognition and text transcription services.

[0026] In step S103, based on the current scenario mode, trigger at least one AI function module. The AI function module includes at least one of an AI auditory module, an AI voice module, an AI visual module, and an AI creation module; When the scene mode is the conference scene mode, the AI auditory module, AI speech module, AI vision module, and AI creation module are triggered; according to requirements, the AI speech module can real-time translate the acquired audio information into audio in different languages and output it synchronously. At the same time, this module will also intelligently adjust the emotional intensity of the translated speech based on non-verbal information such as the gesture amplitude and expression of the keynote speaker captured by the AI vision module to ensure that the speech output is not only accurate but also emotional and closer to the actual communication scenario. Specifically, according to the expression of the keynote speaker obtained by the AI vision module, the speech adjustment actions corresponding to the expression labels are as follows: (1) Label: Confused; Execution: Automatically reduce the speech speed, pop up keyword explanations, and prompt the speaker in the background that further explanations may be needed. (2) Label: Bored / Distracted (wandering eyes); Execution: Prompt the speaker in the background that they may need to switch to the next PPT page and ask if they need to take a break; (3) Label: Approval / Interest (slight smile + nod); Execution: Prompt the speaker in the background to extend the current topic time and highlight the speech content during this period in the meeting minutes; (4) Label: Disapproval / Dissatisfaction (tight lips + shaking head); Execution: Prompt the speaker in the background to ask if there are different opinions; (5) Label: Fatigue (drooping eyelids + yawning); Execution: Prompt the speaker in the background if they need to take a break.

[0027] The AI auditory module plays a key role in the conference. It is responsible for receiving and processing the voice information of the keynote speaker. Since the AI auditory module in the edge device cannot meet the computing power requirements for calling the voiceprint library to identify the identity of the keynote speaker, the pre-stored voiceprint library will be called through the cloud service device to accurately identify the identity of the keynote speaker, effectively distinguish the voices of different participants, so as to accurately judge the identity of the current speaker and the number of participants in the conference.

[0028] The AI vision module focuses on processing the visual information in the conference. It can scan the image information on the shared screen and convert it into text information using technologies such as optical character recognition (OCR) for subsequent collation and recording.

[0029] The AI creation module will exert its powerful integration ability after the conference. It automatically summarizes the information collected and processed by each module during the conference and generates a detailed meeting minutes. Since the AI creation module in the edge device cannot meet the computing power requirements for generating meeting reports, to-do lists, or PPT reports according to the preset templates, the AI model will be called through the cloud service device to obtain meeting reports, to-do lists, or PPT reports, greatly improving the efficiency and convenience of the subsequent work of the conference.

[0030] Obtain the multi-modal information of the user in the current scene mode in step S104, and the multi-modal information includes voice information, audio information, and image information; The voice information mainly comes from the microphone array, which can capture the user's speech and other sounds in the environment, such as the discussion sounds in a meeting or the navigation prompts in a driving scenario. The audio information comes from audio input devices such as microphones or audio interfaces and includes audio signals corresponding to voice commands, background music, etc., such as the explanatory audio in a meeting or the audio playback content in a home environment. The image information is obtained by image acquisition devices such as cameras, presenting the visual environment around the user, including human actions, object shapes, and spatial layouts, such as the expressions and gestures of participants in a meeting scenario, as well as the images on the shared screen, or the road conditions and traffic signs in a driving scenario. The edge device collaborates with multiple AI modules such as hearing, speech, and vision to accurately collect and integrate this multimodal information, laying a foundation for subsequent processing and analysis to achieve a more intelligent and accurate interaction experience.

[0031] In step S105, the multimodal information is preprocessed and the result data that conforms to the current scene mode is output according to a preset rule; The preprocessing includes operations such as speech recognition, semantic analysis, image recognition, and image tagging, converting the voice information into text, extracting the key features of the image, and identifying the key content in the audio. The processed information generates structured result data according to a preset rule. For example, a meeting minutes is generated in a meeting scenario, and device control instructions are generated in a home scenario. The preprocessed information is further processed and optimized, and finally multimodal fusion data is generated as the basis for subsequent decision-making and actions.

[0032] Specifically, in a meeting scenario, the system performs noise reduction processing on the voice information, extracts clear voice signals, and then converts them into text through speech recognition technology. At the same time, key frame extraction and feature recognition are performed on the image information to capture the expressions and gestures of the participants. The preprocessed text and image information are integrated according to a preset rule to generate a meeting minutes, including the meeting theme, speakers, key discussion points, and to-do items, etc. In a home scenario, the system performs semantic analysis on the voice command to identify the user's operation intention, and at the same time performs scene recognition on the image information to determine the home environment where the user is located. The preprocessed information generates device control instructions and sends them to the corresponding smart home devices to achieve intelligent control of the devices.

[0033] In step S106, when the edge device enables the upload permission and the AI function module cannot meet the requirements for multimodal information processing, the result data is sent to the cloud service device, so that the cloud service device processes it according to the priority processing strategy of the multimodal information to generate multimodal fusion data, returns it to the edge device for storage, and decides whether to publish the multimodal fusion data according to the business requirements; synchronize the multimodal fusion data and its publishing information to the mobile terminal device; among them, the priority processing strategy of the multimodal information is determined based on the current scene mode.

[0034] This process ensures the consistency and availability of data across multiple devices, improving the flexibility and efficiency of data processing. Edge devices can autonomously decide on data publication, enhancing the control over data publication, while data synchronization ensures that users can obtain the latest information in a timely manner on different devices. The processing capabilities of cloud service devices make data processing more efficient, reducing the computing burden on edge devices. In this way, the system optimizes resource utilization, improves the efficiency of data processing and publication, while ensuring data consistency and availability, thereby enhancing the overall user experience.

[0035] In the embodiments of this application, when the edge device can meet the computing power requirements for the implementation of each AI function module in the current scenario mode, there is no need to send the result data to the cloud service device for further fusion processing by the AI large model. At this time, the result data is the multi-modal fusion data of the edge device based on the user in the current scenario. The edge device will decide whether to publish the fusion data to each multimedia platform according to the actual business requirements, and will synchronize the fusion data and its publication information to the mobile terminal device at the same time, and the customer on the mobile terminal device side will confirm the fusion data. And in this process, the priority of triggering each AI function by the edge device will be set according to the requirements of real-time performance, security, etc. in the current scenario. For example, in a meeting scenario, the AI function module will be triggered first for speech recognition to ensure that the meeting content can be converted into text in real time. At the same time, the operation of the speech translation function will also be prioritized to enable participants to understand the speech content in a timely manner. In a home scenario, when a security-related instruction (such as "emergency alarm") is detected, the relevant security warning and processing functions will be triggered first to ensure the safety of family members. In a driving scenario, the system will prioritize the processing of information related to driving safety, such as real-time traffic condition analysis and obstacle detection, to help the driver make timely responses. In this way, the edge device can reasonably allocate resources according to the specific requirements of the scenario to ensure the efficient operation of key functions.

[0036] In the embodiments of this application, the edge device establishes a connection with the cloud service device through an adaptive communication protocol. When the edge device is disconnected from the cloud service device, the edge device continues to process the multi-modal information through the locally pre-stored AI function module and automatically synchronizes the data to the cloud service device when reconnected. The edge device uses a local caching algorithm to temporarily store the processed data locally to ensure data is not lost. When the edge device reconnects to the cloud service device, it automatically uploads the locally cached data to the cloud service device through the synchronization mechanism, ensuring the coherence of multi-modal information processing and enhancing the reliability of the system in an unstable network environment.

[0037] In the embodiments of the present application, after the edge device and the mobile terminal device establish a connection for the first time through an adaptive communication protocol, the mobile terminal device will pre-configure the upload permission for the edge device to send the result data to the cloud service device, so that when the upload permission is enabled and the AI function module cannot meet the requirements of multimodal information processing, the edge device will send the result data to the cloud service device. After the upload permission of the edge device is configured once, the edge device will save the permission and automatically upload it to the cloud service device when generating the result data subsequently. The automated upload ensures that the data is timely transmitted to the cloud service device for data fusion processing, which not only guarantees the real-time and accuracy of data processing, but also reduces the need for manual intervention, optimizes the collaborative work of the edge device, the cloud service device and the mobile terminal device, and improves the efficiency of multimodal information processing.

[0038] Among them, the adaptive communication protocol connection includes connection methods such as Wi-Fi 6E / Bluetooth 5.3 / UWB, etc., which can be dynamically selected according to the actual situation. With the design of the adaptive communication protocol and the local AI module, the edge device can still continuously execute the core functions in the offline environment, and automatically synchronize the data to the cloud service device for further processing after reconnecting. This architecture enables the system to complete the tasks in advance even in the weak network environment, avoiding affecting the overall task progress in case of network disconnection or weak network. Embodiment

[0039] Please refer to Figure 2 , an AI multimodal fusion interaction method, applied to a cloud service device, includes: Obtain result data sent by at least one mobile terminal device; the result data is that after the edge device obtains the first voice information of the user, it switches to the corresponding scene mode according to the key information of the first voice information, triggers at least one AI function module based on the current scene mode to obtain the multimodal information of the user in the current scene mode, preprocesses the multimodal information, and outputs the result data that conforms to the current scene mode according to the preset rules; Among them, the key information includes execution action information, scene information or device information; the AI function module includes at least one of an AI auditory module, an AI voice module, an AI visual module, and an AI creation module; the multimodal information includes sound information, audio information and image information; Call the AI large model to perform fusion processing on the result data. The AI large model matches the priority processing strategy of the pre-stored multimodal information according to the scene mode in the result data, processes each modal information in the result data according to the priority processing strategy of the multimodal information, and outputs multimodal fusion data; among them, the priority processing strategy sets the priority according to the scene mode and modal information; Send the multi-modal fusion data to the mobile terminal device so that the mobile terminal device determines whether to publish the multi-modal fusion data.

[0040] After receiving the result data sent by the edge device, the cloud service device calls the AI large model to perform fusion processing on these data. The AI large model matches the pre-stored multi-modal information priority processing strategy according to the scene mode in the result data. For example, in a meeting scene, the model will give priority to processing speech transcription information and extract key discussion points and speaker information; in a driving scene, it will give priority to processing real-time road condition image information to ensure driving safety. According to this strategy, the model processes each modal information in the result data, such as performing semantic analysis on speech information and object recognition on image information, and finally outputs multi-modal fusion data, such as meeting minutes, driving assistance information, etc., and sends this data to the mobile terminal device, and the mobile terminal device decides whether to publish it.

[0041] This process realizes the deep integration and efficient processing of multi-modal information, ensures that key information is processed first, thereby improving the accuracy and timeliness of data processing. With the powerful computing power of the cloud service device, it can perform complex processing on a large amount of data and generate more valuable fusion data. At the same time, after receiving the multi-modal fusion data processed by the cloud service device, the edge device will cache it locally and decide whether to publish the data on the multimedia platform according to the actual business needs, enhancing the flexibility of the system and the user's control over the data. When the AI function module of the edge device cannot meet the requirements of multi-modal information processing, the edge device will send the result data preprocessed and output according to the preset rules to the cloud service device, and the cloud service device will call the AI large model to achieve data processing with complex requirements and high computing power, and realize the fusion processing of multi-modal data. This architecture not only gives full play to the advantages of both, but also optimizes resource allocation and improves the performance and efficiency of the entire system.

[0042] In the embodiments of the present application, according to the priority processing strategy of each piece of modal information, the corresponding modal computing power allocation rules are matched, and each piece of modal information is processed according to the modal computing power allocation rules to output multi-modal fusion data; wherein, the modal computing power allocation rules are set based on the computing power requirements of each piece of modal information in the current scene mode to match the computing power requirements for processing each piece of modal information. For example, in a meeting scene, speech transcription and extraction of key discussion points are given higher priorities. That is, the modal priority processing strategy in the meeting scene is set according to real-time requirements, and the priorities of the corresponding functional modules are: AI speech module > AI auditory module > AI visual module > AI creation module. The device will allocate more computing power to the AI speech module, followed by the AI auditory module, to process speech information quickly and accurately. In a content creation scene, the actions, expressions, and gestures of the creator are given higher priorities, the priority of the AI visual module is higher than that of the AI speech module, and the AI auditory module is used to perform precise sound pickup and sound field analysis first to ensure precise collaboration in single-person and multi-person content creation and high-quality content material collection. By setting the priority processing strategies of each piece of modal information in different scenes, the allocation of computing power resources is optimized, the data processing efficiency is improved, and it is ensured that key information is processed in a timely manner. At the same time, by precisely matching the computing power requirements, resource waste is reduced, the flexibility and adaptability of the system are enhanced, and it can better meet the diverse requirements in different scenes.

[0043] When obtaining multiple pieces of result data sent by multiple edge devices, determine the scene mode of each piece of result data, and process each piece of result data according to the scene computing power allocation rules; wherein, the scene computing power allocation rules are set based on the security requirements, real-time requirements, and general requirements of multiple scene modes to match the computing power requirements of each scene mode.

[0044] Adjust the computing power allocation in different scene modes according to security requirements, real-time requirements, and general requirements. The computing power for security requirements is higher than that for real-time requirements, and the computing power for real-time requirements is higher than that for general requirements. When the security requirements, real-time requirements, and general requirements of the scene are the same, the computing power is evenly divided for processing.

[0045] The scenario computing power allocation rule configuration method is adopted to adjust the computing power of cloud service devices in real time, so as to ensure that when multiple edge devices send processing tasks in various scenario modes to the cloud service devices within the same time period, the cloud service devices can perform real-time allocation of computing power according to security requirements, real-time requirements, and general requirements, ensuring that important and urgent tasks in relevant scenario modes can be completed first. It should be noted that the processing tasks in a certain scenario include multiple subtasks, and each subtask carries corresponding modality information and scenario information. The cloud service device needs to confirm the processing order and computing power allocation rule according to the corresponding modality information and scenario information of each subtask.

[0046] In the driving scenario, the highest priority task is to judge security requirement matters related to life safety. For example, when AEB is triggered, the priority of this matter will be increased within 100ms, all entertainment systems will be forced to stop, and emergency functions related to safety will be enabled, such as issuing instructions to retract the seat belt and unlock the door; while other tasks such as playing music or opening the sunroof are real-time requirements or general requirements. Of course, if the computing power requirements for security requirements or real-time requirements in this scenario are low and the AI function module of the edge device can directly implement function processing, it will be directly processed on the edge device to improve the overall processing efficiency.

[0047] In the home scenario, the highest priority is security requirement matters to ensure safety, such as security alarms and fire alarms are prioritized, followed by all other function controls, such as voice control to turn on or off the air conditioner, automatically adjust the air conditioner, open or close the curtains, and implement energy-saving strategies. Of course, if the computing power requirements for security requirements or real-time requirements in this scenario are low and the AI function module of the edge device can directly implement function processing, it will be directly processed on the edge device to improve the overall processing efficiency.

[0048] In the content creation scenario, the highest priority is to execute the instructions actively issued by the user first, that is, real-time requirements or general requirements. When tasks with security requirements in other scenarios appear, the background is allowed to lower the task priority in the content creation scenario. Of course, if the computing power requirements for security requirements or real-time requirements in this scenario are low and the AI function module of the edge device can directly implement function processing, it will be directly processed on the edge device to improve the overall processing efficiency.

[0049] In a meeting scenario, the highest priority is to ensure real-time synchronization of translated speech to ensure that participants can understand each other's speeches in a timely manner, that is, the real-time requirement; while for real-time synchronization of meeting images, requirements such as displaying translated content through image information and performing image processing on meeting images based on the micro-expressions of participants are general requirements. When there are tasks with security requirements in other scenarios, the background is allowed to lower the task priority in the meeting scenario. Of course, if the computing power requirements for security requirements or real-time requirements in this scenario are low and the AI function module of the edge device can directly implement function processing, the processing will be directly performed on the edge device to improve the overall processing efficiency.

[0050] For example, a certain edge device sends the requirements and related data of a cross-border real-time meeting to the cloud service device, while another edge device sends the requirements and related data in the scenarios of content creation digital humans and driving to the cloud service device. At this time, the cloud service device will first process the security requirement tasks in the driving scenario, then process the real-time requirement tasks in the meeting scenario, content creation scenario, and driving scenario, and finally process the general requirement tasks in the three scenarios. Of course, with continuous use and iteration, the priority tasks in each scenario will be updated and adjusted according to user needs. Moreover, when the security requirements, real-time requirements, or computing power requirements in different scenarios are low, the tasks with security requirements and real-time requirements will be processed first on the edge device according to the priority processing strategy to improve the overall processing efficiency.

[0051] Through the setting and expansion of multi-scenario modes and multi-modal information, the modal information of users in each scenario can be effectively identified to effectively support the preprocessing of edge devices and the large model processing of cloud service devices, ensure the accuracy of information processing and the fusion degree between modal information, and at the same time make the processing order of priorities clearer.

[0052] Furthermore, the functions that the edge device can implement or the scenarios it can support will be informed to the user in advance so that the user can wake up the edge device to execute relevant AI workflows through key information.

[0053] Taking the content creation mode as an example for expansion, it is as follows: When creating a digital human for filming, first, wake up the edge device (such as our company's Yunbao) through voice, such as "I want to create a digital human and publish it". First, Yunbao will use the AI function module to convert the voice into text and capture key information such as creating a digital human. After the edge device recognizes the key information, it will execute, that is, turn on the camera and start voice recording, allowing the shooter to shoot a sample video of 30s - 50s according to the pre-set rules (ensuring that the time is long enough to ensure the sampling effect of AI, and at the same time ensuring the appearance of continuous expressions and sufficient multi-angle facial features). The edge device sends the captured sample video and the pre-prepared copywriting to the cloud service device, and calls the AI model of the cloud service device to generate a realistic digital human video. Using multi-modal task parallelism can save at least half of the time compared to traditional methods, and quickly realize the automated workflow of the entire digital human content creation process.

[0054] When the scene mode is a meeting scene, the audio information, voice information, and image information in the multi-modal information are processed in sequence according to the priority processing strategy, and computing power is allocated to the audio information, voice information, and image information in sequence according to the modal computing power allocation rule.

[0055] In traditional meeting systems, voice input (listening) is converted into text for meeting records, and visual input (reading) is used to identify the speaker and perform camera tracking. This listening and reading are processed independently without interaction, that is, the cross-modal fusion of "listening + reading" is insufficient. It mainly means that when the AI system processes voice input (listening) and text / visual input (reading) simultaneously, it is difficult to establish deep semantic associations, resulting in the inability to accurately understand the multi-modal interaction intentions of users. However, the multi-modal interaction intention solved in this case can be to translate the voice input (listening) into different texts in real-time and overlay them on the screen. Through visual input (reading), the expressions of the participants are recognized. If a frowning expression is found, it will be judged that there is difficulty in understanding, and the translation speed and the overlaid text will be slowed down so that the participants can hear and see the speaker's statement clearly. When the recognized expression returns to normal, the normal speed will be restored. Of course, if it is judged that the entire synchronization speed is too slow for too long, the playback rate will be appropriately increased.

[0056] Based on the multi-modal priority processing strategy adapted to the scene mode, dynamic interaction optimization of cross-modal information is realized. For example, in a meeting scene, the expressions of the keynote speaker are captured through the visual module to real-time correct the emotional intensity and translation rhythm of the voice module; in the digital human creation scene, the expression actions of the virtual image are automatically adjusted in combination with the emotional parameters of the voice command to achieve audio-visual emotional synchronization. This mechanism significantly improves the information understanding accuracy rate of the system in complex scenes, and through the automated workflows preset by the edge device (such as meeting minutes generation, digital human video synthesis), the execution efficiency of multi-modal tasks is increased by 50% - 200% compared to traditional methods. Embodiment

[0057] Reference Figure 4 As shown, an AI-based multi-modal data processing and interaction ecosystem includes: An edge device, which is used to obtain the user's first voice information and identify the key information of the first voice information, and switch to the corresponding scene mode based on the key information. The key information includes execution action information, scene information or device information; Based on the current scene mode, trigger at least one AI function module. The AI function module includes at least one of an AI auditory module, an AI voice module, an AI visual module, and an AI creation module; Obtain the multi-modal information of the user in the current scene mode. The multi-modal information includes sound information, audio information and image information; preprocess the multi-modal information and output result data that conforms to the current scene mode according to a preset rule. When the edge device enables the upload permission and the AI function module cannot meet the multi-modal information processing requirements, send the result data to the cloud service device; A cloud service device, which obtains the result data sent by at least one edge device; Call an AI large model to perform fusion processing on the result data. The AI large model matches the priority processing strategy of the pre-stored multi-modal information according to the scene mode in the result data, and processes each modal information in the result data according to the priority processing strategy of the multi-modal information, and outputs multi-modal fusion data, which is returned to the edge device for storage, so that the edge device decides whether to publish the multi-modal fusion data according to business requirements; wherein, the priority processing strategy sets priorities according to the scene mode and modal information; A mobile terminal device, which receives the multi-modal fusion data and its release information synchronized by the edge device.

[0058] Reference Figure 3 As shown, an AI multi-modal fusion interaction edge device includes: A data recognition module, which is used to obtain the user's first voice information, identify the key information of the first voice information, and switch to the corresponding scene mode based on the key information. The key information includes execution action information, scene information or device information storage unit; An AI function call module, which is used to trigger at least one AI function module based on the current scene mode. The AI function module includes an AI auditory module, an AI voice module, an AI visual module and an AI creation module; A preprocessing module, configured to obtain multimodal information of a user in the current scenario mode, where the multimodal information includes voice information, audio information, and image information; preprocess the multimodal information, and output result data conforming to the current scenario mode according to a preset rule; A communication module, configured to, when the edge device enables the upload permission and the AI function module cannot meet the processing requirements of the multimodal information, send the result data to a cloud service device, so that the cloud service device processes the result data according to the priority processing strategy of the multimodal information to generate multimodal fusion data, and send the multimodal fusion data to a mobile terminal device, so that the mobile terminal device determines whether to publish the multimodal fusion data, returns the multimodal fusion data to the edge device for storage, and decides whether to publish the multimodal fusion data according to service requirements; synchronize the multimodal fusion data and its publication information to the mobile terminal device; where the priority processing strategy of the multimodal information is determined based on the current scenario mode.

[0059] A computing device, the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, where, when the computer program instructions are executed by the processor, the device is triggered to execute any one of the methods described above.

[0060] The method and / or embodiment in the embodiments of the present application can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above functions defined in the method of the present application are executed.

[0061] It should be noted that the computer-readable medium described in the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0062] In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the foregoing.

[0063] The computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0064] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0065] As another aspect, embodiments of the present application further provide a computer-readable medium, which may be included in the device described in the above embodiments; or may exist alone without being assembled into the device. The above computer-readable medium carries one or more computer program instructions, and the computer program instructions can be executed by a processor to implement the steps of the methods and / or technical solutions of the foregoing embodiments of the present application.

[0066] In addition, embodiments of the present application further provide a computer program, which is stored in a computer device, so that the computer device executes the access control method.

[0067] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program (including related data structures) of the present application can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0068] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claimed rights. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims can also be implemented by one unit or device through software or hardware. First, second, etc. are used to denote names and do not denote any particular order.

Claims

1. An AI multimodal fusion interaction method, applied to edge devices, characterized in that: include: Obtaining the user's first voice information; identifying key information of the first voice information, and switching to a corresponding scene mode based on the key information; Based on the current scene mode, trigger at least one AI function module to obtain multimodal information of the user in the current scene mode; Preprocessing the multimodal information and outputting result data that conforms to the current scene mode according to preset rules; When the edge device opens the upload permission and the AI ​​function module cannot meet the multimodal information processing requirements, the result data is sent to the cloud service device, so that the cloud service device processes the multimodal information according to the priority processing strategy of the multimodal information to generate multimodal fusion data, returns it to the edge device for storage, and decides whether to publish the multimodal fusion data according to business needs; Synchronizing the multimodal fusion data and its release information to a mobile terminal device; Among them, the AI ​​functional module includes at least one of an AI auditory module, an AI voice module, an AI visual module, and an AI creation module; the multimodal information includes sound information, audio information, and image information; wherein the priority processing strategy of the multimodal information is determined based on the current scene mode.

2. The AI ​​multimodal fusion interaction method according to claim 1, characterized in that: include: The edge device establishes a connection with the cloud service device through an adaptive communication protocol. When the edge device is disconnected from the cloud service device, the edge device continues to process the multimodal information through a locally pre-stored AI function module and automatically synchronizes data to the cloud service device when reconnected.

3. The AI ​​multimodal fusion interaction method according to claim 1, characterized in that: When the edge device enables upload permission and the AI ​​function module cannot meet the multimodal information processing requirements, the result data is sent to the cloud service device, including: after the edge device establishes a connection with the mobile terminal device, the mobile terminal device will pre-configure the upload permission of the edge device to send the result data to the cloud service device, so that the edge device sends the result data to the cloud service device when the upload permission is enabled.

4. The AI ​​multimodal fusion interaction method according to claim 1, characterized in that: include: When the scene mode is a conference scene mode, triggering the AI ​​hearing module, AI voice module, AI vision module and AI creation module based on the current conference scene mode; The AI ​​voice module synchronously outputs the translated audio of the audio information according to the demand, and adjusts the voice emotion intensity according to the speaker's gesture amplitude and expression captured by the AI ​​vision module; The AI ​​hearing module obtains the voice information of the speaker; The AI ​​vision module scans the image information in the shared screen and generates text information; The AI ​​creation module automatically summarizes the meeting minutes after the meeting.

5. An AI multimodal fusion interaction method, applied to a cloud service device, characterized in that: include: Obtain result data sent by at least one edge device; The result data is that after the edge device obtains the first voice information of the user, it switches to the corresponding scene mode according to the key information of the first voice information, and triggers at least one AI function module based on the current scene mode to obtain the multimodal information of the user in the current scene mode, pre-processes the multimodal information, and outputs the result data that conforms to the current scene mode according to the preset rules; The AI ​​function module includes at least one of an AI hearing module, an AI voice module, an AI vision module, and an AI creation module; the multimodal information includes sound information, audio information, and image information; Calling the AI ​​big model to perform fusion processing on the result data, the AI ​​big model matches the pre-stored priority processing strategy of the multimodal information according to the scene mode in the result data, processes each modal information in the result data according to the priority processing strategy of the multimodal information, and outputs multimodal fusion data; wherein the priority processing strategy sets the priority according to the scene mode and modal information; The multimodal fusion data is returned to the edge device for storage, so that the edge device decides whether to publish the multimodal fusion data according to business needs, and synchronizes the multimodal fusion data and its publishing information to the mobile terminal device.

6. The AI ​​multimodal fusion interaction method according to claim 5, characterized in that: Processing each modality information in the result data according to the priority processing strategy of the multimodal information to output multimodal fusion data includes: According to the priority processing strategy of each modal information, the corresponding modal computing power allocation rule is matched, each modal information is processed according to the modal computing power allocation rule, and multimodal fusion data is output; wherein the modal computing power allocation rule is set based on the computing power requirement of each modal information under the current scene mode to match the computing power requirement for processing each modal information.

7. The AI ​​multimodal fusion interaction method according to claim 5, characterized in that: Calling the AI ​​big model to perform fusion processing on the result data includes: When multiple result data sent by multiple edge devices are obtained, the scene mode of each result data is determined, and each result data is processed according to the scene computing power allocation rule; wherein the scene computing power allocation rule is set based on the security requirements, real-time requirements and general requirements of multiple scene modes to match the computing power requirements of each scene mode.

8. The AI ​​multimodal fusion interaction method according to claim 6, characterized in that: When the scene mode is a conference scene, the audio information, sound information and image information in the multimodal information are processed in sequence according to the priority processing strategy, and computing power is allocated to the audio information, sound information and image information in sequence according to the modal computing power allocation rule.

9. An AI multimodal fusion interactive edge device, characterized in that: include: A data recognition module, used to obtain first voice information of a user, identify key information of the first voice information, and switch to a corresponding scene mode based on the key information; An AI function calling module, used to trigger at least one AI function module based on the current scene mode, wherein the AI ​​function modules include an AI hearing module, an AI voice module, an AI vision module and an AI creation module; A preprocessing module, used to obtain multimodal information of the user in the current scene mode, wherein the multimodal information includes sound information, audio information and image information; Preprocessing the multimodal information and outputting result data that conforms to the current scene mode according to preset rules; A communication module, which is used to send the result data to the cloud service device when the edge device opens the upload permission and the AI ​​function module cannot meet the multimodal information processing requirements, so that the cloud service device processes the multimodal information according to the priority processing strategy of the multimodal information to generate multimodal fusion data, returns it to the edge device for storage, and decides whether to publish the multimodal fusion data according to business needs; The multimodal fusion data and its release information are synchronized to a mobile terminal device; wherein the priority processing strategy of the multimodal information is determined based on the current scene mode.

10. A multimodal fusion interaction system based on AI, characterized in that: include: an edge device, the edge device being used to obtain first voice information of a user and identify key information of the first voice information, and switch to a corresponding scene mode based on the key information; Based on the current scene mode, at least one AI function module is triggered, wherein the AI ​​function module includes at least one of an AI hearing module, an AI voice module, an AI vision module, and an AI creation module; Acquire multimodal information of the user in the current scene mode, wherein the multimodal information includes sound information, audio information and image information; Preprocessing the multimodal information and outputting result data that conforms to the current scene mode according to preset rules; When the edge device opens the upload permission and the AI ​​function module cannot meet the multimodal information processing requirements, the result data is sent to the cloud service device; The cloud service device obtains result data sent by at least one of the edge devices; Calling the AI ​​big model to perform fusion processing on the result data, the AI ​​big model matches the pre-stored priority processing strategy of the multimodal information according to the scene mode in the result data, processes each modal information in the result data according to the priority processing strategy of the multimodal information, outputs the multimodal fusion data, and returns it to the edge device for storage, so that the edge device decides whether to publish the multimodal fusion data according to business needs; wherein the priority processing strategy sets the priority according to the scene mode and modal information; The mobile terminal device receives the multimodal fusion data and its publishing information synchronized by the edge device.

11. A computing device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein: When the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Intelligent interactive remote auxiliary video system

    CN113965550A

  • Multi-modal fusion sensing architecture and method for multiple mobile robots

    CN117873100A

  • Multi-modal information pushing method, storage medium and electronic device

    CN118612001A

  • Multi-mode large-model edge accelerator system and implementation method, device, equipment and medium thereof

    CN119106717A

  • Vehicle information interaction method and system and vehicle

    CN119682681A

Cited By

  • AI workflow processing method and system of intelligent hardware

    CN120547367A

  • An AI workflow processing method and system for intelligent hardware

    CN120547367B

  • Intelligent conference recording and recording method and system

    CN120568005A

  • Intelligent conference recording and recording method and system

    CN120568005B

  • End-side multi-mode fusion analysis and active service interaction method and system

    CN121580285A