Multi-modal data processing method, model deployment method, equipment and storage medium

By deploying large models on mobile devices, processing multimodal data and filtering the information users are concerned about based on semantic similarity, privacy leakage and response speed problems caused by cloud deployment are solved, and personalized services and user experience are improved.

CN120086555APending Publication Date: 2025-06-03BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510240033.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The deployment of existing large-scale models in the cloud has led to a decrease in user privacy leakage risks and response speeds. At the same time, users hope that the large-scale models can provide personalized services based on user behavior, but how to intelligently identify user behavior and provide customized services has become a challenge.

Method used

Deploy a large model on a mobile device, calculate the semantic similarity between the data specified by multimodal data (such as images, sounds, locations, etc.), and use the large model to process data when meeting the preset similarity threshold, filter out the information that users are concerned about and provide personalized services.

Benefits of technology

By processing multimodal data on the mobile side, data upload to the cloud is avoided, the risk of privacy leakage is reduced, the use of mobile computing power is reduced, and the correlation between user experience and large-scale model calculation results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086555A_ABST
    Figure CN120086555A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data processing method, a model deployment method, equipment and a storage medium, the method is applied to a mobile terminal, the mobile terminal is deployed with a large model, and the method comprises the steps of obtaining multi-modal data related to a target user; calculating the semantic similarity between the multi-modal data and specified data of the large model; and when the semantic similarity meets a preset similarity threshold range, calculating a data processing result corresponding to the multi-modal data by using the large model. According to the method and the device, the data leakage risk can be reduced, the information concerned by the user can be screened out, calculation is carried out for the information concerned by the user, the data volume of model processing is reduced, occupation of mobile terminal computing power is reduced, the association degree between a large model calculation result and the information concerned by the user can be improved, and the mobile terminal use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a multimodal data processing method, a model deployment method, a device, and a storage medium. Background Art

[0002] Large models have been widely used in various fields. Due to their large number of parameters, they are mostly deployed on high-computing-power GPU (Graphics Processing Unit) devices and cloud servers. With the support of high computing power, large models have a fast inference speed, but they also bring cost and privacy issues. Moreover, as the functions of large models become more abundant, users are no longer satisfied with simply chatting with large models, and also hope that large models can perform more intelligent interactions based on user behavior. Summary of the Invention

[0003] Embodiments of this application provide a multimodal data processing method, a model deployment method, a device, and a storage medium to solve one or more of the above technical problems.

[0004] In a first aspect, an embodiment of this application provides a multimodal data processing method. The method is applied to a mobile device on which a large model is deployed. The method includes: obtaining multimodal data related to a target user; calculating the semantic similarity between the multimodal data and specified data of the large model respectively; and when the semantic similarity meets a preset similarity threshold range, using the large model to calculate a data processing result corresponding to the multimodal data.

[0005] In a second aspect, an embodiment of this application provides a model deployment method, including: obtaining a first model and a second model, where the number of parameters of the first model is greater than that of the second model; using the first model to calculate a first output result corresponding to target input data, and using the second model to calculate a second output result corresponding to the target input data; calculating a loss using the first output result and the second output result, adjusting the second model according to the loss, and deploying the adjusted second model on the mobile device to use the mobile device to execute the method according to any one of claims 1-6.

[0006] In a third aspect, an embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method according to any one of the above is implemented.

[0007] In a fourth aspect, an embodiment of this application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method according to any one of the above is implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method described in any one of the above.

[0009] Compared with the related art, the present application has the following advantages:

[0010] The present application provides a multi-modal data processing method, a model deployment method, a device, and a storage medium. The method is applied to a mobile device on which a large model is deployed. The method includes: obtaining multi-modal data related to a target user; calculating the semantic similarity between the multi-modal data and specified data of the large model respectively; when the semantic similarity meets a preset similarity threshold range, using the large model to calculate the data processing result corresponding to the multi-modal data. According to the embodiments of the present application, the large model deployed on the mobile device can be used to process multi-modal data, thereby avoiding uploading relevant multi-modal data to the cloud and reducing the risk of data leakage; calculating the semantic similarity between the multi-modal data and the specified data of the large model respectively, and by calculating the semantic similarity, when the semantic similarity meets the preset similarity threshold range, using the large model to calculate the data processing result corresponding to the multi-modal data, the information concerned by the user can be screened out, the calculation is performed on the information concerned by the user, the amount of data processed by the model is reduced, the occupancy of the computing power of the mobile device is reduced. In addition, by calculating the semantic similarity and screening out the information concerned by the user, the correlation between the calculation result of the large model and the information concerned by the user can also be improved, and the user experience of using the mobile device can be improved.

[0011] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0013] Figure 1 Shows a flowchart of the multi-modal data processing method provided in the embodiment of the present application;

[0014] Figure 2 Shows a data processing flowchart of a knowledge distillation provided in the embodiment of the present application;

[0015] Figure 3Shows a schematic diagram of the engineering deployment of a large model provided in an embodiment of the present application;

[0016] Figure 4 Shows a schematic diagram of an intelligent life interaction assistant provided in an embodiment of the present application;

[0017] Figure 5 Shows a schematic diagram of a multi-round information processing mechanism provided in an embodiment of the present application;

[0018] Figure 6 Shows a flowchart of an intelligent recognition interaction provided in an embodiment of the present application;

[0019] Figure 7 Shows a flowchart of data preprocessing provided in an embodiment of the present application;

[0020] Figure 8 Shows a flowchart of personalized service provided in an embodiment of the present application;

[0021] Figure 9 Shows a flowchart of a model deployment method provided in an embodiment of the present application;

[0022] Figure 10 Shows a structural block diagram of a multi-modal data processing device provided in an embodiment of the present application;

[0023] Figure 11 Shows a structural block diagram of a model deployment device provided in an embodiment of the present application;

[0024] Figure 12 Shows a block diagram of an electronic device for implementing the embodiments of the present application. Detailed implementation manners

[0025] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature and not restrictive.

[0026] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.

[0027] In the current digital age, with the rapid development of artificial intelligence technology, large deep learning models (referred to as large models) have demonstrated unprecedented capabilities in multiple fields such as natural language processing, image recognition, and recommendation systems. These models, trained with massive amounts of data, can perform complex tasks such as understanding natural language, generating text, and analyzing images, greatly enhancing the intelligent level of human-computer interaction. However, traditionally, these large models are often deployed on cloud servers, and users need to interact with them via the Internet. Although this model is flexible and easy to maintain, there are risks of user privacy leakage in some scenarios, and the response speed may also be affected by network conditions. As people's attention to personal privacy increases, users are increasingly reluctant to upload sensitive information to the cloud for processing. Moreover, with the enrichment of the functions of large models, users are no longer satisfied with simply having conversations with large models, and also hope that large models can provide personalized services based on user behavior. Therefore, how to intelligently identify user behavior and provide customized services to users has become an urgent problem to be solved.

[0028] Based on this, the present application provides a multi-modal data processing method, a model deployment method, a device, and a storage medium. By deploying a large model on a mobile device, intelligent dialogue interaction with users is realized. Then, various modal data of users, including but not limited to information such as images, sounds, and locations, are collected by collecting home intelligent devices or wearable intelligent devices, etc. Combining with large model technology, various behaviors in users' daily lives are intelligently identified and analyzed, and personalized services are provided according to the recognition results.

[0029] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the foregoing technical problems. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0030] The embodiments of the present application provide a multi-modal data processing method. This method is applied to a mobile terminal, and a large model is deployed on the mobile terminal. Among them, a large model (Large Model) is a deep learning model with a large number of parameters and a complex structure in the fields of artificial intelligence and machine learning. A mobile terminal is an application program and service running on a mobile device. Mobile devices usually include portable devices such as smart phones, tablets, and smart watches. These devices usually have wireless connection functions and can access the Internet and various services anytime and anywhere. As Figure 1 shown in the flowchart of the multi-modal data processing method of an embodiment of the present application, this method may include:

[0031] Step S101, obtain multi-modal data related to a target user.

[0032] In the embodiments of the present application, a user is an individual or entity using a certain system, platform, application or service, which may include but are not limited to: individual users, enterprise users, registered users and anonymous users. The target user can be a fixed user. By obtaining multimodal data related to the target user, it is convenient to provide customized services for a fixed user.

[0033] In the embodiments of the present application, multimodal data is data containing multiple types or forms. These data can come from different perceptual channels or sensors and reflect different aspects of the same object or phenomenon. The integration and analysis of multimodal data can provide a more comprehensive and in-depth understanding, especially in application fields that require integrating multiple information sources. Multimodal data includes but is not limited to the following types of data: image data, video data, audio data, text data, sensor data, structured data, geospatial data, and time series data.

[0034] It should be noted that multimodal data can be data generated by different devices. For example, it can be data generated by smart home devices or wearable smart devices. These devices for generating multimodal data can be communicatively connected to the mobile device where the mobile terminal is located.

[0035] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0036] Step S102, calculate the semantic similarity between the multimodal data and the specified data of the large model respectively.

[0037] In the embodiments of the present application, the specified data of the large model is data pre-determined to be related to the large model, including but not limited to: data that has been processed by the large model and data that is expected to be processed by the large model.

[0038] The specified data of the large model can include multiple groups of data. For each group of data, the semantic similarity with the multimodal data can be calculated, that is, calculate the semantic similarity between the multimodal data and the specified data of the large model respectively.

[0039] In this step, by calculating the semantic similarity, data that the user is interested in can be mined from the multimodal data, so as to perform targeted analysis and calculation on the data that the user is interested in. For example, multimodal data with a higher semantic similarity to the specified data of the large model can be used as the data that the user is more interested in.

[0040] Step S103, when the semantic similarity meets the preset similarity threshold range, use the large model to calculate the data processing result corresponding to the multimodal data.

[0041] In the embodiments of the present application, the preset similarity threshold range can be set according to actual needs. For example, it can be set to be greater than a certain threshold, less than a certain threshold, greater than or equal to a certain threshold, less than or equal to a certain threshold, between two thresholds, etc. The specific setting method is not specifically limited in this application.

[0042] In this step, the data processing result can be the data output by the large model. When the semantic similarity meets the preset similarity threshold range, the multimodal data that the user is interested in can be screened out, and the multimodal data that the user is interested in can be combined with the large model for calculation, so as to obtain the data processing result corresponding to the multimodal data, reduce the amount of calculation, improve the calculation efficiency, and make the data processing result closer to the user's attention.

[0043] The present application provides a multimodal data processing method, a model deployment method, a device, and a storage medium. This method is applied to a mobile device on which a large model is deployed. The method includes: obtaining multimodal data related to a target user; calculating the semantic similarity between the multimodal data and the specified data of the large model respectively; when the semantic similarity meets the preset similarity threshold range, using the large model to calculate the data processing result corresponding to the multimodal data. According to the embodiments of the present application, the large model deployed on the mobile device can be used to process multimodal data, thereby avoiding uploading relevant multimodal data to the cloud and reducing the risk of data leakage; calculating the semantic similarity between the multimodal data and the specified data of the large model respectively. By calculating the semantic similarity, when the semantic similarity meets the preset similarity threshold range, using the large model to calculate the data processing result corresponding to the multimodal data can screen out the information that the user is interested in, perform calculations on the information that the user is interested in, reduce the amount of data processed by the model, reduce the occupancy of the computing power of the mobile device. In addition, by calculating the semantic similarity and screening out the information that the user is interested in, the correlation between the calculation result of the large model and the information that the user is interested in can also be improved, and the user experience of using the mobile device can be improved.

[0044] Considering the limitations of mobile devices in computing power, directly supporting the processing of overly long contexts poses a significant challenge. An overly long context may lead to issues such as mobile device memory explosion, affecting the user experience. In contrast, the cloud environment, with its powerful computing resources, can easily accommodate longer context lengths, support the introduction of more rounds of conversation history, and thus guide the large model to make more detailed and comprehensive responses. Mobile devices, limited by their finite computing power, struggle to handle such high-demand tasks. Therefore, in one possible implementation, when the specified data is the historical conversation query data of the large model and the semantic similarity meets the preset similarity threshold range, the data processing result corresponding to the multimodal data can be calculated using the large model as follows: When the semantic similarity exceeds the preset similarity threshold, sort the historical conversation query data according to the semantic similarity, combine the sorted result and the multimodal data to obtain the first input data, and calculate the data processing result corresponding to the first input data using the large model; when the semantic similarity does not exceed the preset similarity threshold, use the multimodal data as the second input data and calculate the data processing result corresponding to the second input data using the large model.

[0045] In this possible implementation, the historical conversation query data of the large model refers to the conversation records and query data accumulated by the large model (such as ChatGPT) during multiple rounds of interaction with the user. When the semantic similarity exceeds the preset similarity threshold, it indicates a high degree of correlation between the multimodal data and the historical conversation query data of the large model. Sort the historical conversation query data according to the semantic similarity, and combine the sorted result and the multimodal data, that is, concatenate the historical conversation query data with the user's current query data in the order of the sorted result to obtain the first input data.

[0046] When the semantic similarity does not exceed the preset similarity threshold, that is, the correlation between the multimodal data and the historical conversation query data of the large model is low. In this case, the historical conversation query data cannot play a positive role and may even affect the large model's answer. At the same time, an overly long context input will also increase the resource occupancy of the mobile device. Therefore, use the multimodal data as the second input data, that is, no longer concatenate it with the historical conversation query data, use the second input data as the input of the large model, and calculate the data processing result corresponding to the second input data using the large model, thereby improving the large model's computing efficiency and reducing the occupancy of computing resources.

[0047] For the multi-round conversation scenario, the history of the previous conversations is concatenated for each query. Suppose the user asks four questions in a row as follows:

[0048] Introduce the famous tourist attractions in Beijing, what's your name, introduce large model technology, and what to have for lunch.

[0049] These four questions have insufficient relevance to each other and lack a problem theme. If the large model has answered the first three questions and the fourth question needs to be asked, then the first three conversation histories and the fourth query need to be concatenated as the input to the large model. Obviously, the fourth question has no relation to the first three questions, and its historical conversation information cannot play a positive role, and may even affect the answer of the large model. At the same time, too long context input will also increase the resource occupancy of mobile devices. See Figure 5 The schematic diagram of the multi-round information processing mechanism shown in the figure. In the figure, Query1, Query2, ……, Queryn represent historical conversation query data, Query represents the user's current query data, and Response represents the processing result of the large model. Based on the embodiments of the present application, the following steps can be executed:

[0050] (1) First, when compiling the model, it is necessary to consider the performance of the deployment device to compile a suitable input length model, and this length is used as the upper limit of the input length of the large model.

[0051] (2) Store the session history on the mobile device. Respectively judge the semantic similarity between the current query and the queries in the historical conversation to obtain a similarity value.

[0052] (3) Set a threshold to judge whether the similarity value between the current conversation query and the historical conversation query exceeds the threshold. If it exceeds the threshold, then sort the queries and answers corresponding to the historical conversation from large to small according to the similarity, and loop to judge whether the length of the concatenated result exceeds the input length upper limit. If it exceeds, discard the historical conversation with small similarity. If it does not exceed, directly input it into the large model to generate an answer; if it does not exceed the threshold, it means that the current query theme has no relation to the historical conversation, then directly input the current query into the large model to generate an answer.

[0053] In a complex and changeable interaction scenario, it is crucial to retain appropriate multi-round conversation history. In this possible implementation, it can effectively assist the large model to more deeply understand the user's intention and provide more accurate and demand-oriented answers.

[0054] In a possible implementation, combine the sorting result and the multi-modal data to obtain the first input data. The following steps can be executed: Combine the sorting result and the multi-modal data to obtain a combined result; if the length of the combined result exceeds the input length upper limit of the large model, then adjust the combined result according to the sorting result until the length of the obtained combined result does not exceed the input length upper limit to obtain the first input data.

[0055] In this possible implementation, the sorting result and the multimodal data are combined to obtain a combined result. Specifically, for example, the sorting result includes the order SDEAM, where each letter represents a historical conversation query data, and each historical conversation query data includes query data and the answer data of the large model for this query data. The user's current query data is B, and the current query data includes the current query data. Concatenating the historical conversation query data with the user's current query data can obtain SDEAMB or BSDEAM, which is the combined result.

[0056] The upper limit of the input length of the large model is predetermined. For example, when compiling the model, it is necessary to consider the performance of the deployment device to compile a suitable input length model. If the length of the combined result exceeds the upper limit of the input length of the large model, the combined result is adjusted according to the sorting result. For example, if the upper limit of the input length of the large model is limited to a fixed-length character, then it is necessary to reduce the length of the sorting result in the combined result, and then determine whether the length of the newly obtained combined result does not exceed the input length upper limit. If it still exceeds, continue to delete a historical conversation query data from the sorting result until the length of the obtained combined result does not exceed the input length upper limit, and the first input data is obtained. Then, the first input data is used as the input of the large model, and the data processing result corresponding to the first input data is calculated by using the large model, so that the large model can give a data processing result that better meets the user's needs in combination with the historical conversation query data.

[0057] In this possible implementation, for application scenarios with insufficient mobile computing power and multi-round conversation requirements, multi-round information processing can be performed to retain an appropriate length for the first input data, thereby optimizing the answer of the large model and reducing the resource occupancy of the mobile device.

[0058] In one possible implementation, adjusting the combined result according to the sorting result can be performed according to the following steps: determining the historical conversation query data with the smallest semantic similarity in the current sorting result; deleting the historical conversation query data with the smallest semantic similarity in the combined result.

[0059] In this possible implementation, for example, for the combined result SDEAMB, the current sorting result is SDEAM, where each letter represents a historical conversation query data, and the multiple historical conversation query data are arranged in descending order according to the value of semantic similarity. Determine the historical conversation query data with the smallest semantic similarity in the current sorting result, that is, M, and delete the historical conversation query data with the smallest semantic similarity in the combined result, that is, M needs to be deleted. Then, determine whether the length of SDEAB does not exceed the input length upper limit. If it still exceeds, delete A.

[0060] Users have many behaviors in their daily lives, such as eating, sleeping, physical exercise, taking medicine, etc. These behaviors may vary due to personal habits and other reasons, resulting in different levels of attention from users. Users may not need the large model to remind them of some behaviors. In a possible implementation manner, when the specified data is pre-configured behavior data and the semantic similarity meets the preset similarity threshold range, the large model is used to calculate the data processing result corresponding to the multimodal data, and the following steps can be executed: when the semantic similarity exceeds the preset similarity threshold, the large model is used to calculate the reminder information corresponding to the multimodal data; the reminder information includes the attribute data corresponding to the behavior data; and the reminder information is used as the data processing result.

[0061] In this possible implementation manner, the pre-configured behavior data can be obtained by pre-custom configuration and is used to represent the behaviors that users are concerned about. For example, it includes but is not limited to the following behaviors: applying a facial mask, taking medicine, and exercising, etc. When the semantic similarity meets the preset similarity threshold range, it indicates that the similarity between the multimodal data and the pre-configured behavior data is relatively high. Therefore, the multimodal data is used as the input of the large model, and the large model is used to calculate the reminder information corresponding to the multimodal data, so as to remind the user with the reminder information. Among them, the reminder information includes the attribute data corresponding to the behavior data, and the attribute data is used to represent what response behavior is expected from the large model. The type of the attribute data can be predefined, and for different behavior data, different attribute data can be included.

[0062] Taking applying a facial mask as an example, the pre-configured behavior data corresponding to applying a facial mask is as follows. Among them, "concerned behavior" and "what response behavior is expected from the large model" are the attribute labels of the behavior data and are mandatory items, and the content inside can be customized by the user.

[0063] Concerned behavior: Applying a facial mask;

[0064] What response behavior is expected from the large model: including but not limited to recording the start time of applying a facial mask, giving scientific advice on the duration of applying a facial mask, and intelligently reminding the user whether the best time for applying a facial mask has been reached.

[0065] Among them, the above-mentioned facial mask application time data is obtained by devices such as home intelligent devices or wearable intelligent devices, and all these data are used as the input of the large model for the large model to analyze and judge. When the large model recognizes that the user is applying a facial mask, it judges according to the custom behavior library that this behavior is a behavior that the user is concerned about, and then gives a reminder notification according to "what response behavior is expected from the large model" in the behavior library.

[0066] In this possible implementation, a large model deployed on a mobile device is used to analyze the processed data, and corresponding services are triggered based on the recognition results. This application supports users to customize the behavior library that the large model needs to focus on. The large model intelligently determines whether to send reminder messages by recognizing user behaviors and the customized behavior library.

[0067] Figure 8 A personalized service flow chart is shown. In a possible implementation, the data with noise processed is input into the large model, and the large model automatically recognizes these data. By calculating the semantic similarity between the recognition results and the user-defined behavior library, it is determined whether there are behaviors that the user is concerned about. If not, no processing is performed; if so, combined with the user's preferences and historical behavior data, reminder messages for the user are intelligently generated and sent to the user's mobile device in the form of system notifications.

[0068] In a possible implementation, calculating the semantic similarity between the multimodal data and the specified data of the large model can be performed according to the following steps: calculating the semantic similarity between the multimodal data and the specified data of the large model through a string matching method.

[0069] In this possible implementation, the string matching method can be used to calculate the similarity between texts. The semantic similarity between the multimodal data and the specified data of the large model can be calculated by calculating the edit distance (measuring the difference between two strings, common algorithms include Levenshtein distance, Damerau-Levenshtein distance, etc.), Jaccard similarity (calculating the ratio of the intersection and union of two sets. For strings, n-gram (such as bigram or trigram) can be used to convert strings into sets), Cosine similarity (representing strings as vectors and then calculating the cosine similarity between vectors), or Jaro-Winkler similarity (a measure for calculating the similarity between two strings, especially suitable for short strings).

[0070] In this possible implementation, in scenarios where short texts are processed or quick similarity calculation is required, the string matching method can be used to calculate text similarity. The embodiments of this application provide a simple and effective means to measure the similarity between texts.

[0071] See Figure 4 As shown in the schematic diagram of an intelligent life interaction assistant, the implementation of this method will be described below with a specific example.

[0072] Based on this method, the intelligent interaction assistant function can be realized. Specifically, by using the large model deployed on the mobile device and the collected user behavior data, an intelligent interactive life assistant based on the large model is designed. This assistant includes two major modules: an intelligent recognition and interaction module and an intelligent dialogue module, as Figure 4 shown. Intelligent dialogue module: This module uses the large model to conduct intelligent conversations with users. Since the mobile device may have insufficient computing power, a multi-round information processing mechanism is designed to avoid the problem of the mobile device crashing due to overly long context information. Intelligent recognition and interaction module: This module collects image, sound, location information, etc. through smart home devices or wearable smart devices, analyzes the user's behavior using the large model, and automatically generates system reminder information based on the recognition results.

[0073] Among them, the intelligent recognition and interaction module integrates advanced image recognition, sound recognition, and location information acquisition technologies. By collecting the user's image, sound, and location data in daily life, it deeply analyzes the user's behavior pattern using the large model, intelligently recognizes the user's current state and potential needs, and automatically generates and pushes personalized system reminder information accordingly, aiming to improve the user experience and enhance the convenience and security of life. See Figure 6 the intelligent recognition and interaction flowchart shown. First, collect data containing user behavior, in forms including but not limited to images, sounds, location information, etc. These data can be obtained through the user's smart home devices or wearable smart devices, and are transmitted to the user's mobile device through wireless communication technologies (such as Wi-Fi), mobile applications (such as smart home APPs), etc. The user can set the data transmission time interval according to actual needs, and then the large model carried by the mobile device is used for automatic recognition and analysis, and corresponding personalized services are provided according to the recognition results. The specific implementation details are as follows:

[0074] The data information transmitted from the smart device to the mobile device may contain noise or redundant information. If directly processed by the large model, it will result in long processing time, and at the same time, the noise information will also affect the large model effect. Therefore, data preprocessing is required first. The data most likely to have noise are images and sounds, and these two types of data are easily affected by the environment. In this application, traditional denoising algorithms are used to process these two types of data, as Figure 7 shown. The specific approach is as follows:

[0075] For image data, the mean filtering denoising algorithm is used. By taking the mean of pixel points and then assigning the calculated mean to the current pixel point as the gray value of the processed image at that point. The formula is:

[0076]

[0077] Among them, f(x, y) represents the restored image, g(r, c) represents the noisy image, and s xy represents the coordinates of a rectangular sub-window with center (x, y) and size m×n.

[0078] The voice data adopts an adaptive filtering algorithm. By using the results of the filter parameters obtained at the previous moment, it automatically adjusts the filter parameters at the current moment to adapt to the unknown or time-varying statistical characteristics of the signal and noise, thereby achieving optimal filtering. An adaptive filter is essentially a Wiener filter that can adjust its own transmission characteristics to achieve optimality. The adaptive filter does not require prior knowledge of the input signal and has a small computational load, making it particularly suitable for real-time processing.

[0079] As Figure 7 shown, x(n) represents the input signal of the adaptive filter, usually the signal to be processed or the signal containing noise, d(n) is the target output signal, usually the ideal signal expected to be obtained. It is used to calculate the error signal to guide the adjustment of the filter. y(n) is the estimated output of the adaptive filter for the input signal x(n) under the current weights. e(n) is obtained by processing the input signal through the filter. It is the difference between the desired signal d(n) and the filter output y(n), that is, e(n) = d(n) - y(n). The error signal reflects the deviation between the filter output and the desired signal, and is the basis for the adaptive algorithm to adjust the filter weights. Weight update (system update) is the process of adjusting the filter weights according to the error signal e(n). The adaptive algorithm uses the error signal to update the filter weights to gradually reduce the error signal and make the filter output closer to the desired signal. Through the interaction of these input and output signals, the adaptive filter can dynamically adjust its parameters to optimize the filtering effect and minimize the output error.

[0080] In the embodiments of the present application, an intelligent assistant function is provided. The intelligent assistant can not only conduct intelligent conversations, but also customize user attention actions, perform intelligent recognition and analysis, and give reminder notifications.

[0081] The present application provides a multi-modal data processing method, a model deployment method, a device, and a storage medium. By deploying a large model on a mobile device, it can ensure that user data is processed locally, reduce the risk of data leakage, and enhance users' confidence in privacy protection. For application scenarios with insufficient mobile computing power and multi-round conversation requirements, the present application provides a multi-round information processing mechanism to manage the problem of overly long context generated by multi-round conversation history. In addition, the present application can also collect various modal data of users, use the large model to identify and analyze the data, and give corresponding notification services according to the user-defined behavior library.

[0082] Given the natural limitations of mobile devices in computing power, directly running large models with tens of billions or even tens of billions of parameters (such as models with 6 billion to 14 billion parameters) on mobile devices is particularly challenging. Therefore, in practice, it is more inclined to select models with relatively small parameter scales (such as 500 million to 1.8 billion parameters) to adapt to the mobile platform environment. However, it cannot be ignored that the expansion of model scale is often accompanied by a significant increase in the number of parameters, which not only enhances the memory capacity of the model, enabling it to capture complex features in data more comprehensively, but also directly promotes a leap in the accuracy and quality of the generated results. In contrast, small-scale models often struggle to match the comprehensiveness of data capture and the depth of generated content. If such models are deployed on mobile devices, it will inevitably sacrifice the richness and satisfaction of the user experience to a certain extent. Therefore, while pursuing mobile application efficiency and user experience, how to balance model scale and performance has become an important topic in current technical exploration and optimization.

[0083] Therefore, the embodiments of this application also provide a model deployment method, as Figure 9 shown in the flowchart of the model deployment method according to an embodiment of this application. The method may include:

[0084] Step S901, obtain a first model and a second model; wherein, the number of parameters of the first model is greater than the number of parameters of the second model.

[0085] In the embodiments of this application, the first model or the second model may be a large-scale deep learning model used in natural language processing (NLP) and generation tasks. For example, it includes but is not limited to the following models: GPT (Generative Pre-trained Transformer) series models including GPT-2, GPT-3, GPT-4, etc.; BERT (Bidirectional Encoder Representations from Transformers) model. The number of parameters of the two is different, wherein the number of parameters of the first model is greater than the number of parameters of the second model.

[0086] Step S902, use the first model to calculate a first output result corresponding to the target input data, and use the second model to calculate a second output result corresponding to the target input data.

[0087] In the embodiment of the present application, according to the insufficient computing power of the mobile terminal, the knowledge distillation technology is used to distill the large-scale parameter model into a model with smaller parameters, so as to improve the accuracy of the small-scale parameter model without increasing the parameters. The large-scale model (first model) is regarded as the teacher model, and the small model (second model) is regarded as the student model. The student model imitates the teacher model, and the two compete with each other, so that the student model can perform on par with or even better than the teacher model.

[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0089] Step S903: Calculate the loss using the first output result and the second output result, adjust the second model according to the loss, and deploy the adjusted second model on the mobile terminal to execute any of the above methods using the mobile terminal.

[0090] In the embodiments of the present application, Figure 2 The data processing flow chart of a knowledge distillation is shown. The knowledge adopts relationship-based knowledge to further explore the relationship between different data samples, and uses multiple outputs to combine into structural units, which can better reflect the structured characteristics of the teacher model and enable the student model to be better guided.

[0091] The loss function for relational distillation learning ( Figure 2 Distillation Loss) is as follows, where T1, T2…Tn represent the teacher model ( Figure 2 Teacher) multiple outputs, S1, S2...Sn represent student models ( Figure 2 In the figure, there are multiple outputs of Student), and L represents the calculation of the distance between the two.

[0092]

[0093] The student model that has undergone the above distillation has achieved significant improvements in performance compared to the original small-scale model. Not only is it closer to the excellent teacher model in terms of effect, but also, most importantly, its number of parameters has not increased at all, which makes it an ideal choice for mobile deployment. Given the natural limitations of mobile devices in terms of performance and resources, it is obviously impractical to run the original complex model directly on the device. Therefore, efficiently compiling and converting the model to adapt to the mobile environment has become a crucial step.

[0094] This application can adopt the MLC-LLM (Multilingual Large Language Model) framework to carry out the engineering compilation work of the model. This framework can be applied to multiple platforms, easily handle the compilation requirements of different mobile devices, and greatly simplify the user's integration process. With the support of the MLC-LLM framework, the student model can be seamlessly connected to various mobile devices, retaining its excellent performance advantages while taking into account the resource limitations of the mobile end, facilitating the development, optimization, and deployment of models on multiple types of mobile devices, and bringing users an unprecedented convenient and efficient experience.

[0095] See Figure 3 As shown in the schematic diagram of the engineering deployment of a large model, the distilled student model is quantized. In the figure, Model definition represents model definition, Model compilation represents model compilation, Platform-native runtimes represents platform-native runtimes, and APK represents the file format for distributing and installing Android applications. Whether quantization is required and the quantization bit number if quantization is required can be selected according to the mobile device to be deployed and the requirements. The student model is compiled using the MLC-LLM framework, and finally the compiled model is deployed on the mobile end to use the mobile end to execute any one of the above-mentioned multimodal data processing methods.

[0096] This application provides a model deployment method, which can deploy the adjusted small model on the mobile end and use the mobile end to implement the multimodal data processing method. The multimodal data processing method includes: obtaining multimodal data related to the target user; calculating the semantic similarity between the multimodal data and the specified data of the large model respectively; when the semantic similarity meets the preset similarity threshold range, using the large model to calculate the data processing result corresponding to the multimodal data. According to the embodiments of this application, the large model deployed on the mobile end can be used to process multimodal data, thereby avoiding uploading relevant multimodal data to the cloud and reducing the risk of data leakage; calculating the semantic similarity between the multimodal data and the specified data of the large model respectively, and by calculating the semantic similarity, when the semantic similarity meets the preset similarity threshold range, using the large model to calculate the data processing result corresponding to the multimodal data, the information concerned by the user can be screened out, calculate for the information concerned by the user, reduce the amount of data processed by the model, reduce the occupancy of the computing power of the mobile end. In addition, by calculating the semantic similarity and screening out the information concerned by the user, the correlation between the calculation result of the large model and the information concerned by the user can also be improved, and the mobile end usage experience of the user can be improved.

[0097] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application further provide a multimodal data processing device. As Figure 10 shown in the structural block diagram of the multimodal data processing device according to an embodiment of the present application, the device may include:

[0098] A multimodal module 1001, configured to obtain multimodal data related to a target user; a similarity module 1002, configured to calculate the semantic similarity between the multimodal data and the specified data of the large model respectively; a processing module 1003, configured to, when the semantic similarity meets a preset similarity threshold range, use the large model to calculate the data processing result corresponding to the multimodal data.

[0099] In a possible implementation manner, when the specified data is the historical dialogue query data of the large model, the processing module is specifically configured to: when the semantic similarity exceeds a preset similarity threshold, sort the historical dialogue query data according to the semantic similarity, combine the sorting result and the multimodal data to obtain first input data, and use the large model to calculate the data processing result corresponding to the first input data; when the semantic similarity does not exceed the preset similarity threshold, use the multimodal data as second input data, and use the large model to calculate the data processing result corresponding to the second input data.

[0100] In a possible implementation manner, combining the sorting result and the multimodal data to obtain first input data includes: combining the sorting result and the multimodal data to obtain a combined result; if the length of the combined result exceeds the input length upper limit of the large model, adjust the combined result according to the sorting result until the length of the obtained combined result does not exceed the input length upper limit, so as to obtain first input data.

[0101] In a possible implementation manner, adjusting the combined result according to the sorting result includes: determining the historical dialogue query data with the smallest semantic similarity in the current sorting result; deleting the historical dialogue query data with the smallest semantic similarity in the combined result.

[0102] In a possible implementation manner, when the specified data is pre-configured behavior data, the processing module is specifically configured to: when the semantic similarity exceeds a preset similarity threshold, use the large model to calculate a reminder message corresponding to the multimodal data; the reminder message includes attribute data corresponding to the behavior data; use the reminder message as the data processing result.

[0103] In a possible implementation, the similarity module is specifically configured to: calculate the semantic similarity between the multimodal data and the specified data of the large model through a string matching method.

[0104] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application also provide a model deployment device. As Figure 11 shown in the structural block diagram of the model deployment device according to an embodiment of the present application, the device may include:

[0105] A model module 1101, configured to obtain a first model and a second model; wherein, the number of parameters of the first model is greater than that of the second model; a calculation module 1102, configured to calculate a first output result corresponding to the target input data by using the first model, and calculate a second output result corresponding to the target input data by using the second model; a deployment module 1103, configured to calculate a loss by using the first output result and the second output result, adjust the second model according to the loss, and deploy the adjusted second model on the mobile device to execute the multimodal data processing method described in any one of the above.

[0106] For the functions of each module in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and the corresponding beneficial effects are achieved, which will not be elaborated here.

[0107] Figure 12 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 12 shown, the electronic device includes: a memory 1201 and a processor 1202, and a computer program that can run on the processor 1202 is stored in the memory 1201. When the processor 1202 executes the computer program, the method in the above embodiments is implemented. The number of the memory 1201 and the processor 1202 may be one or more.

[0108] The electronic device further includes:

[0109] A communication interface 1203, configured to communicate with external devices and perform data interaction and transmission.

[0110] If the memory 1201, the processor 1202, and the communication interface 1203 are implemented independently, the memory 1201, the processor 1202, and the communication interface 1203 can be interconnected through a bus and communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 it is represented by only one thick line in Figure 12 , but this does not mean that there is only one bus or one type of bus.

[0111] Optionally, in specific implementation, if the memory 1201, the processor 1202, and the communication interface 1203 are integrated on a chip, the memory 1201, the processor 1202, and the communication interface 1203 can communicate with each other through an internal interface.

[0112] An embodiment of the present application provides a computer-readable storage medium that stores a computer program, and when the program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0113] An embodiment of the present application provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0114] An embodiment of the present application further provides a chip, which includes a processor for calling and running an instruction stored in a memory from the memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.

[0115] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory, and when the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0116] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0117] Further, optionally, the above-mentioned memory can include a read-only memory and a random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can include a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can include a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Sync link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0118] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0119] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0120] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0121] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.

[0122] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices.

[0123] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0124] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a disk, an optical disc, etc.

[0125] As mentioned above, the above are only exemplary embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art in the technical scope recorded in the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multimodal data processing method, wherein: The method is applied to a mobile terminal, where a large model is deployed, and the method includes: Acquire multimodal data related to target users; Calculating semantic similarities between the multimodal data and the designated data of the large model respectively; When the semantic similarity satisfies a preset similarity threshold range, the large model is used to calculate a data processing result corresponding to the multimodal data.

2. The method according to claim 1, wherein: In the case where the designated data is historical conversation query data of the large model, when the semantic similarity satisfies a preset similarity threshold range, the large model is used to calculate a data processing result corresponding to the multimodal data, including: When the semantic similarity exceeds a preset similarity threshold, sorting the historical conversation query data according to the semantic similarity, combining the sorting result and the multimodal data to obtain first input data, and calculating a data processing result corresponding to the first input data using the large model; When the semantic similarity does not exceed the preset similarity threshold, the multimodal data is used as second input data, and the data processing result corresponding to the second input data is calculated using the large model.

3. The method according to claim 2, wherein: The sorting result and the multimodal data are combined to obtain first input data, including: Combining the sorting result and the multimodal data to obtain a combined result; If the length of the combination result exceeds the upper limit of the input length of the large model, the combination result is adjusted according to the sorting result until the length of the obtained combination result does not exceed the upper limit of the input length, thereby obtaining the first input data.

4. The method according to claim 3, wherein: Adjusting the combination result according to the sorting result includes: Determine the historical conversation query data with the smallest semantic similarity in the current ranking result; The historical conversation query data with the smallest semantic similarity is deleted from the combined result.

5. The method according to claim 1, wherein: In the case where the designated data is pre-configured behavior data, when the semantic similarity satisfies a preset similarity threshold range, the data processing result corresponding to the multimodal data is calculated using the large model, including: When the semantic similarity exceeds a preset similarity threshold, the large model is used to calculate the reminder information corresponding to the multimodal data; the reminder information includes the attribute data corresponding to the behavior data; The reminder information is used as the data processing result.

6. The method according to claim 1, wherein: Calculating the semantic similarity between the multimodal data and the designated data of the large model respectively, comprising: The semantic similarity between the multimodal data and the designated data of the large model is calculated by a string matching method.

7. A model deployment method, comprising: Acquire a first model and a second model; wherein the parameter amount of the first model is greater than the parameter amount of the second model; Calculating a first output result corresponding to the target input data using the first model, and calculating a second output result corresponding to the target input data using the second model; Calculate the loss using the first output result and the second output result, adjust the second model according to the loss, and deploy the adjusted second model on the mobile terminal to execute the method described in any one of claims 1-6 using the mobile terminal.

8. An electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer program product, wherein: The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.