A method for a multi-modal cloud fusion service platform

By using API interface to collect data, polynomial kernel functions and deep neural network models for fusion analysis in multimodal data processing, the problem of lack of effective fusion methods in traditional multimodal data processing is solved, and efficient integration of multimodal data and personalized service provision are achieved.

CN119538184BActive Publication Date: 2025-05-30HUAIAN RUIMING INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411554326.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-05-30
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Traditional multimodal data processing methods are usually mainly single-modal processing, and the lack of effective multimodal fusion methods leads to the inability to fully exert the correlation and complementarity between different modal data, affecting the effectiveness and accuracy of data processing, and failing to meet the diverse needs of users, making it difficult to provide personalized services.

Method used

Data from different modes are collected through the API interface, preprocessing and feature extraction, and data fusion and analysis are used using polynomial kernel functions and deep neural network models. A multimodal cloud fusion service platform is provided to provide personalized services according to user needs and application scenarios.

Benefits of technology

Effectively integrate different modal data, give full play to the complementarity of multimodal data, improve the accuracy of analysis and classification, meet the diverse needs of users, provide personalized multimodal services, and improve user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538184B_ABST
    Figure CN119538184B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for a multi-modal cloud fusion service platform, which relates to the technical field of cloud service platforms and includes the following steps: S1. Use an API interface to dock with the user's data source to collect different modal data, and obtain the user's fusion service requirements and application scenarios; S2. Preprocess the collected data, establish a cloud storage database to provide multi-modal cloud storage services, and establish data indexes, etc. The present invention collects different modal data through various methods, expands the data source, improves the richness and diversity of the data, improves the data quality through data preprocessing, provides a reliable data basis for subsequent analysis, and various data index methods facilitate the rapid retrieval and query of data, improving the efficiency of data storage and management. By establishing user portraits and clustering analysis, it is possible to better understand user needs, provide personalized multi-modal services for different user groups, and improve user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud service platforms, and specifically to a method for a multi-modal cloud fusion service platform. Background Art

[0002] With the rapid development of information technology, various sensors, devices, and network platforms continuously generate a large amount of multi-modal data, including text, images, audio, video, etc. Social media platforms are filled with rich graphic, audio, and video content; Internet of Things devices continuously collect various environmental data and device status data. The development of cloud computing technology provides powerful computing and storage resources for the multi-modal cloud fusion service platform. The cloud computing platform can provide elastic computing capabilities, dynamically allocate computing resources according to user needs, and meet the high computing requirements for multi-modal data processing and analysis.

[0003] Currently, traditional multi-modal data processing methods usually focus on single-modal processing and lack effective multi-modal fusion methods. This fails to fully utilize the correlation and complementarity between data of different modalities, affecting the effect and accuracy of data processing. Moreover, they usually can only provide a single service, unable to meet the diverse needs of users and difficult to provide personalized services according to user needs and preferences. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for a multi-modal cloud fusion service platform to solve the problems in the above background art that traditional multi-modal data processing methods usually focus on single-modal processing and lack effective multi-modal fusion methods. This fails to fully utilize the correlation and complementarity between data of different modalities, affecting the effect and accuracy of data processing. Moreover, they usually can only provide a single service, unable to meet the diverse needs of users and difficult to provide personalized services according to user needs and preferences.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for a multi-modal cloud fusion service platform, including the following steps:

[0006] S1. Use the API interface to connect to the user's data source to collect different modalities of data, and obtain the user's fusion service requirements and application scenarios. The API interface and specific data collection technologies ensure the efficiency and accuracy of data collection;

[0007] S2. Preprocess the collected data, establish a cloud storage database to provide multi-modal cloud storage services and establish data indexes. The data indexes include inverted indexes, spatial indexes, and hybrid indexes. The API interface and specific data collection technologies ensure the efficiency and accuracy of data collection;

[0008] S3. Extract features for data of different modalities, and represent the extracted features in a unified way. The unified representation can better reflect the comprehensive information of multimodal data, facilitating subsequent fusion and analysis.

[0009] S4. Based on the method of polynomial kernel function, fuse the extracted data features. Use a deep neural network model to automatically learn the associations and fusion patterns between different modalities, evaluate and optimize the fusion results. The feature extraction methods for different modality data can fully explore the characteristics and values of each modality data. The fusion method based on polynomial kernel function and the deep neural network model can effectively integrate different modality data, give play to the complementarity of multimodal data, and improve the accuracy of analysis and classification.

[0010] S5. Provide multimodal services according to user requirements and application scenarios, and deploy the multimodal cloud fusion service platform to the cloud.

[0011] S6. Provide a user interaction interface for users to upload multimodal data, query service results, and manage personal information.

[0012] S7. Regularly maintain and update the platform, monitor the running status of the platform, and collect user feedback. Regular maintenance, update, and monitoring of the platform running status ensure the stability and performance of the platform. Collecting user feedback helps to continuously improve the platform functions and services, and enhance the adaptability and competitiveness of the platform.

[0013] Preferably, in step S1, the API interface uses JSON or XML as the data exchange format to provide interactions between different systems. In the multimodal data collection, for image data, use API collection and web crawler technology to collect relevant images from specific image websites, and use object detection algorithms based on deep learning to automatically identify and classify image content; for audio data, use direct collection or speech recognition technology to convert audio into text.

[0014] Preferably, in step S1, the obtaining of user fusion service requirements and application scenarios includes the following steps:

[0015] S11. Obtain user behavior data, preference data, and requirement data, integrate the user data collected from different channels and data sources, establish a unified user data warehouse and perform data cleaning, and perform feature extraction and analysis on the cleaned data.

[0016] S12. Construct a user portrait model, use the K-Means clustering method to group users. For each user data point in the dataset, calculate its distance from each cluster center, and according to the principle of the closest distance, assign the data point to the corresponding cluster, dividing users into different groups.

[0017] S13. Provide personalized services for different groups, including personalized recommendations, customized functional services, and exclusive content creation and services.

[0018] Preferably, in step S2, the data preprocessing includes text data cleaning and format conversion processing, image data enhancement processing, and data missing value processing. The missing value processing uses the K-nearest neighbor algorithm to fill in the missing values, which specifically includes the following steps:

[0019] S21. For two sample points A = (a 1 , a 2 , …, a n ) and B = (b 1 , b 2 , …, b n ), the Euclidean distance calculation formula is:

[0020]

[0021] S22. Calculate the distance between the filling sample X and other samples, and select the k samples with the closest distance as the nearest neighbors. Let the values of the corresponding features of these k nearest neighbor samples be v 1 , v 2 , …, v n ;

[0022] S23. Assign different weights according to the sample distance. The closer the sample distance, the greater the weight. Let the distance between sample X and the nearest neighbor sample i be d(X, i), and the weight w i is defined as where ε is a very small positive number. After determining the k nearest neighbor samples and their corresponding weights w 1 , w 2 , …, w k , as well as the values of the corresponding features of the nearest neighbor samples v 1 , v 2 , …, v k , the filling value is calculated in the way of weighted average, and the formula is as follows:

[0023]

[0024] Preferably, in step S3, the feature extraction includes the following steps:

[0025] S31. Text data: Use the RNN recurrent neural network model to model the text sequence and extract semantic features;

[0026] S32. Image data: Use the convolutional neural network to highlight the important regions in the image for feature extraction;

[0027] S33. Audio Data: Mel Frequency Cepstral Coefficients (MFCCs) are combined with a Convolutional Neural Network (CNN) to extract audio features. The specific process is as follows:

[0028] ① Audio preprocessing: The continuous audio signal is segmented into several short audio frames, each frame being 20 - 40 milliseconds in length, with a certain overlap between frames to ensure the continuity of the audio signal. A window function is applied to each audio frame for windowing, and the audio signal is pre-emphasized. A first-order high-pass filter is used to enhance the high-frequency part of the audio signal, compensate for the attenuation of the speech signal in the high-frequency band, and highlight the high-frequency formants of the audio signal;

[0029] ② The windowed audio frames are subjected to a Fast Fourier Transform (FFT) to convert the time-domain signal into a frequency-domain signal, obtaining a spectrum. The spectrum is passed through a Mel filter bank to convert the linear frequency scale to the Mel frequency scale. The Mel filter bank is a set of triangular band-pass filters evenly distributed on the Mel frequency scale. The conversion relationship between the Mel frequency and the actual frequency is:

[0030]

[0031] where f is the actual frequency. The logarithmic energy of the output of each Mel filter is taken to obtain the log Mel spectrum. The log Mel spectrum is subjected to a Discrete Cosine Transform (DCT) to obtain the Mel Frequency Cepstral Coefficients (MFCCs). The first 13 - 20 coefficients are taken as the feature representation of the audio;

[0032] ③ The calculated MFCC features are used as the input to the Convolutional Neural Network (CNN) to extract audio features and convert them into a fixed-length feature vector.

[0033] Preferably, in step S4, the method based on the polynomial kernel function for fusing the extracted data features includes the following steps:

[0034] S41. For each modality of data, calculate its kernel matrix. The data of the two modalities are respectively set as X = (x 1 , x 2 , …, x n ) and Y = (y 1 , y 2 , …, y m );

[0035] where the kernel matrix of X is K X ; K X (i, j) = K(x i , x j ), and the kernel matrix of Y is K Y ; K Y (i, j) = K(y i , y j );

[0036] S42. Weightedly fuse the kernel matrices of different modalities:

[0037] K f = αK X + (1 - α)K Y ;

[0038] where K f is the fused kernel matrix, and α is the weight parameter;

[0039] S43. Based on the fused kernel matrix, use the support vector machine SVM to perform classification in the kernel space.

[0040] Preferably, in step S4, the process of evaluating and optimizing the fusion result is as follows: According to the specific application scenario, define specific evaluation metrics, then divide the dataset into k equal parts, use k - 1 parts for training each time, and use the remaining one part for validation. Repeat this process k times, record the evaluation results of each fold and calculate the average value, count various evaluation metrics, analyze the performance of the model in different categories or different modalities, identify the advantages and disadvantages of the model, analyze the confusion matrix, find the categories where the model is confused, and identify the directions for improvement.

[0041] Preferably, in step S5, the platform is deployed to the cloud using containerized deployment technology, and the continuous integration and continuous deployment processes are implemented to ensure that the updates and improvements of the platform are synchronized and go online.

[0042] Preferably, in step S6, when the user uploads multi-modal data, a visual data upload progress bar is provided to let the user understand the upload status and progress, and the uploaded data is previewed and verified in real time to ensure that the data quality and format meet the requirements; in the query service result, the collaborative filtering algorithm is used to provide query results that meet the user's needs according to the user's historical queries and preferences.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] 1. In the present invention, different modalities of data are collected in various ways, expanding the data source, improving the richness and diversity of the data, improving the data quality through data preprocessing, providing a reliable data basis for subsequent analysis, facilitating fast retrieval and query of data through various data indexing methods, improving the efficiency of data storage and management, establishing user portraits and clustering analysis, being able to better understand user needs, providing personalized multi-modal services for different user groups, and improving user satisfaction.

[0045] 2. In the present invention, through the fusion method based on polynomial kernel function and the deep neural network model, different modality data can be effectively integrated, the complementarity of multi-modal data can be utilized, the accuracy of analysis and classification can be improved, and the fusion result can be evaluated and optimized. By defining evaluation metrics, cross-validation, analyzing evaluation results and confusion matrices, etc., the advantages and disadvantages of the model and the improvement direction can be identified.

[0046] 3. In the present invention, multi-modal services are provided according to user requirements and application scenarios. The platform is deployed to the cloud using containerization deployment technology, and continuous integration and continuous deployment processes are implemented to ensure synchronous updates. By providing a user interaction interface, including a visual upload progress bar, real-time preview, and verification of uploaded data, a collaborative filtering algorithm is used to provide query results that meet user requirements and a function for managing personal information. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flowchart of a method for a multi-modal cloud fusion service platform of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] Please refer to Figure 1 , the present invention provides a technical solution: a method for a multi-modal cloud fusion service platform, including the following steps:

[0050] Step 1: Use the API interface to connect to the user's data source to collect different modality data, and obtain the user's fusion service requirements and application scenarios. The API interface and specific data collection technologies ensure the efficiency and accuracy of data collection. At the same time, in order to improve the security of the interface, authentication mechanisms such as API keys and OAuth can be adopted to prevent unauthorized access, thereby effectively protecting the security of user data;

[0051] Among them, the API interface uses JSON or XML as the data exchange format to provide interaction between different systems. In multi-modal data collection, for image data, API collection and web crawler technologies are used to collect relevant images from specific image websites, and object detection algorithms based on deep learning are used to automatically identify and classify image content; for audio data, direct collection or speech recognition technology is used to convert audio into text;

[0052] Among them, obtaining the user's fusion service requirements and application scenarios includes the following steps:

[0053] 11) Obtain user behavior data, preference data, and demand data, integrate the user data collected from different channels and data sources, establish a unified user data warehouse and perform data cleaning, and conduct feature extraction and analysis on the cleaned data;

[0054] 12) Construct a user portrait model, use the K-Means clustering method to group users. For each user data point in the dataset, calculate its distance from each cluster center, and according to the principle of the closest distance, assign this data point to the corresponding cluster, dividing users into different groups;

[0055] 13) Provide personalized services for different groups, including personalized recommendations, customized function services, and exclusive content creation and services, which can better understand user needs, provide personalized multi-modal services for different user groups, and improve user satisfaction;

[0056] Step 2: Preprocess the collected data, establish a cloud storage database to provide multi-modal cloud storage services and establish data indexes. The data indexes include inverted indexes, spatial indexes, and hybrid indexes. Through the multi-source data index method, it is convenient to quickly retrieve and query data, improving the efficiency of data storage and management;

[0057] Establish a cloud storage database to provide multi-modal cloud storage services. The cloud storage database has advantages such as high reliability and high scalability, and can meet the storage needs of multi-modal data. At the same time, establish data indexes, including inverted indexes, spatial indexes, and hybrid indexes. The inverted index is suitable for text data and can quickly retrieve text content according to keywords. The spatial index is suitable for data with spatial features such as image data and can retrieve according to spatial positions. The hybrid index combines the advantages of the inverted index and the spatial index and is suitable for the retrieval of multi-modal data. Through various data index methods, it is convenient to quickly retrieve and query data, improving the efficiency of data storage and management;

[0058] Among them, data preprocessing includes text data cleaning and format conversion processing, image data enhancement processing, and data missing value processing. The missing value processing uses the K-nearest neighbor algorithm to fill in the missing values, which can effectively fill in the missing values in the data and improve the integrity of the data. The specific steps are as follows:

[0059] 21) For two sample points A = (a 1 , a 2 , …, a n ) and B = (b 1 , b 2 , …, b n ), the Euclidean distance calculation formula is:

[0060]

[0061] 22) Calculate the distances between the filled sample X and other samples, and select the k samples with the closest distances as the nearest neighbors. Let the values of the corresponding features of these k nearest neighbor samples be v 1 , v 2 , …, v n ;

[0062] 23) Assign different weights according to the sample distances. The closer the sample distance is, the greater the weight. Let the distance between sample X and the nearest neighbor sample i be d(X, i), and the weight w i is defined as where ε is a very small positive number. After determining the k nearest neighbor samples and their corresponding weights w 1 , w 2 , …, w k , as well as the values of the corresponding features of the nearest neighbor samples v 1 , v 2 , …, v k are determined, the filled value is calculated in the way of weighted average, and the formula is as follows:

[0063]

[0064] Step 3: Extract features for different modalities of data, and represent the extracted features uniformly. Uniform representation can better reflect the comprehensive information of multi-modal data and facilitate subsequent fusion and analysis; feature extraction includes the following steps:

[0065] 31) Text data: Use the RNN recurrent neural network model to model the text sequence and extract semantic features;

[0066] 32) Image data: Use the convolutional neural network to highlight the important regions in the image for feature extraction;

[0067] 33) Audio data: Use the Mel-frequency cepstral coefficients combined with the convolutional neural network to extract audio features. The specific process is as follows:

[0068] ① Audio preprocessing: Segment the continuous audio signal into several short audio frames. The length of each frame is 20 - 40 milliseconds, and there is a certain overlap between frames to ensure the continuity of the audio signal. Apply a window function to each audio frame for windowing, perform pre-emphasis processing on the audio signal, and enhance the high-frequency part of the audio signal through a first-order high-pass filter to compensate for the attenuation of the speech signal in the high-frequency band and highlight the high-frequency resonance peak of the audio signal;

[0069] ②Perform a fast Fourier transform on the windowed audio frames to convert the time-domain signal into a frequency-domain signal, obtaining a spectrum. Pass the spectrum through a Mel filter bank to convert the linear frequency scale to the Mel frequency scale. The Mel filter bank is a set of triangular band-pass filters evenly distributed on the Mel frequency scale. The conversion relationship between the Mel frequency and the actual frequency is as follows:

[0070]

[0071] where f is the actual frequency. Take the logarithmic energy of the output of each Mel filter to obtain the log Mel spectrum. Perform a discrete cosine transform on the log Mel spectrum to obtain the Mel frequency cepstral coefficients. Take the first 13 - 20 coefficients as the feature representation of the audio;

[0072] ③Use the calculated MFCC features as the input of the convolutional neural network to extract audio features and convert them into fixed-length feature vectors;

[0073] Step Four: Use a method based on polynomial kernel function to fuse the extracted data features. Utilize a deep neural network model to automatically learn the associations and fusion patterns between different modalities, evaluate and optimize the fusion results. The feature extraction methods for different modality data can fully exploit the characteristics and values of each modality data. The fusion method based on polynomial kernel function and the deep neural network model can effectively integrate different modality data, give play to the complementarity of multi-modal data, and improve the accuracy of analysis and classification;

[0074] Among them, the method based on polynomial kernel function to fuse the extracted data features includes the following steps:

[0075] 41) For the data of each modality, calculate its kernel matrix. The data of the two modalities are respectively set as X = (x 1 , x 2 , …, x n ) and Y = (y 1 , y 2 , …, y m );

[0076] Among them, the kernel matrix K X of X; K X (i, j) = K(x i , x j ), and the kernel matrix K Y of Y; K Y (i, j) = K(y i , y j ). By calculating the kernel matrix, the data of different modalities can be mapped into a high-dimensional space, making the data of different modalities have better separability in the high-dimensional space;

[0077] 42) Perform weighted fusion on kernel matrices of different modalities:

[0078] K f = αK X + (1 - α)K Y ;

[0079] where K f is the fused kernel matrix, α is the weight parameter. By adjusting the weight parameter, the contribution degree of different modalities in the fusion can be controlled, and the advantages of different modalities can be fully utilized to improve the accuracy of the fusion result;

[0080] 43) Based on the fused kernel matrix, use the support vector machine SVM to perform classification in the kernel space, which can make full use of the fusion features of multi-modal data and improve the accuracy of classification;

[0081] Specifically, the process of evaluating and optimizing the fusion result is as follows: According to the specific application scenario, define specific evaluation indicators, then divide the data set into k equal parts. Each time, k - 1 parts are used for training and the remaining one part is used for verification. Repeat this k times, record the evaluation results of each fold and calculate the average value, count various evaluation indicators, analyze the performance of the model under different categories or different modalities, identify the advantages and disadvantages of the model, analyze the confusion matrix, find the categories where the model is confused and identify the directions for improvement; Through this cross-validation method, more stable and reliable evaluation results can be obtained. Count various evaluation indicators, analyze the performance of the model under different categories or different modalities, identify the advantages and disadvantages of the model. For the deficiencies of the model, the parameters of the kernel function can be adjusted, the structure of the deep neural network can be optimized, the weight parameter can be adjusted, etc. to improve the accuracy and stability of the fusion result;

[0082] Step Five: Provide multi-modal services according to user requirements and application scenarios, and deploy the multi-modal cloud fusion service platform to the cloud;

[0083] Among them, the platform is deployed to the cloud using containerization deployment technology, and continuous integration and continuous deployment processes are implemented to ensure that the updates and improvements of the platform are synchronized online. Containerization deployment and continuous integration and continuous deployment processes ensure the high availability, scalability and rapid update of the platform, adapting to the changing user requirements and technological development;

[0084] Step Six: Provide a user interface for users to upload multi-modal data, query service results and manage personal information;

[0085] Among them, the user uploads multi-modal data to provide a visual data upload progress bar, enabling the user to understand the upload status and progress, and conduct real-time preview and verification of the uploaded data to ensure that the data quality and format meet the requirements; in querying service results, the collaborative filtering algorithm is used to provide query results that meet the user's needs based on the user's historical queries and preferences. By analyzing the user's historical behavior and the behavior of other users, user groups similar to the current user are identified, and then relevant service results are recommended for the current user based on the behavior of these similar users, improving the accuracy and personalization of the user's query results and meeting the user's needs;

[0086] Step 7: Regularly maintain and update the platform, monitor the running status of the platform, and collect user feedback. Regular maintenance, update, and monitoring of the platform running status ensure the stability and performance of the platform, and collecting user feedback helps continuously improve the platform functions and services, enhancing the adaptability and competitiveness of the platform.

[0087] In this method, different modal data is collected through various means, expanding the data sources, improving the richness and diversity of the data. Establishing user portraits and clustering analysis can better understand user needs, provide personalized multi-modal services for different user groups, and improve user satisfaction. Data preprocessing improves data quality, providing a reliable data basis for subsequent analysis. Multiple data indexing methods facilitate fast retrieval and query of data, improving the efficiency of data storage and management. The fusion method based on polynomial kernel function and the deep neural network model can effectively integrate different modal data, leveraging the complementarity of multi-modal data and improving the accuracy of analysis and classification. Containerized deployment and continuous integration and continuous deployment processes ensure the high availability, scalability, and rapid update of the platform, adapting to changing user needs and technological developments, providing a good user experience, and facilitating users to upload data, query results, and manage personal information.

[0088] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for a multimodal cloud fusion service platform, characterized in that: The following steps are involved: S1. Use the API interface to connect to the user's data source to collect data of different modalities, and obtain the user's fusion service requirements and application scenarios; S2. Preprocess the collected data, establish a cloud storage database to provide multimodal cloud storage services and establish data indexes, which include inverted indexes, spatial indexes and hybrid indexes; S3, extracting features from data of different modalities and expressing the extracted features in a unified manner; S4, based on the polynomial kernel function method, the extracted data features are fused, the deep neural network model is used to automatically learn the association and fusion mode between different modalities, and the fusion results are evaluated and optimized; S5. Provide multimodal services according to user needs and application scenarios, and deploy the multimodal cloud fusion service platform to the cloud; S6. Provide a user interaction interface for users to upload multimodal data, query service results and manage personal information; S7. Regularly maintain and update the platform, monitor the platform's operating status, and collect user feedback; In step S2, the data preprocessing includes text data cleaning and format conversion processing, image data enhancement processing and data missing value processing. The missing value processing uses the K nearest neighbor algorithm to fill in the missing values ​​and specifically includes the following steps: S21, for two sample points A=(a1, a2,…, a n ) and B=(b1,b2,…,b n ), the Euclidean distance calculation formula is: S22. Calculate the distance between the filled sample X and other samples, select the k samples with the closest distance as the nearest neighbors, and set the corresponding feature values ​​of these k nearest neighbor samples to be v1, v2, …, v n ; S23. Different weights are assigned according to the sample distance. The closer the distance, the greater the weight. Suppose the distance between sample X and its nearest neighbor sample i is d(X,i), and the weight w i Defined as Where ε is a very small positive number. After determining the k nearest neighbor samples and their corresponding weights w1, w2, …, w k , and the values ​​of the corresponding features of the neighboring samples v1,v2,…,v k After that, fill in the value Calculated in the weighted average way, the formula is as follows: In step S4, the method based on the polynomial kernel function to fuse the extracted data features includes the following steps: S41. For each modal data, calculate its kernel matrix. The data of the two modalities are set as X = (x1, x2, ..., x n ) and Y=(y1,y2,…,y m ); Among them, the kernel matrix K of X X ; K X (i,j)=K(x i ,x j ), the kernel matrix K of Y Y ; K Y (i,j)=K(y i ,y j ); S42. Weighted fusion of kernel matrices of different modes: K f =αK X +(1-α)K Y ; Among them, K f is the fused kernel matrix, α is the weight parameter; S43. Based on the fused kernel matrix, use support vector machine (SVM) to perform classification in the kernel space.

2. The method of a multimodal cloud fusion service platform according to claim 1, characterized in that: In step S1, the API interface uses JSON or XML as a data exchange format to provide interaction between different systems; in the multimodal data collection, for image data, API collection and web crawler technology are used to collect relevant images from a specific image website, and a deep learning-based target detection algorithm is used to automatically identify and classify image content; for audio data, direct collection or speech recognition technology is used to convert audio into text.

3. The method of a multimodal cloud fusion service platform according to claim 1, characterized in that: In step S1, obtaining user fusion service requirements and application scenarios includes the following steps: S11. Obtain user behavior data, preference data and demand data, integrate user data collected from different channels and data sources, establish a unified user data warehouse and perform data cleaning, and perform feature extraction and analysis on the cleaned data; S12. Build a user portrait model and use the K-Means clustering method to group users. For each user data point in the data set, calculate its distance to each cluster center. According to the principle of the closest distance, assign the data point to the corresponding cluster and divide the users into different groups. S13. Provide personalized services for different groups, including personalized recommendations, customized functional services, and exclusive content creation and services.

4. The method of a multimodal cloud fusion service platform according to claim 1, characterized in that: In step S3, the feature extraction includes the following steps: S31, text data: Use the RNN recurrent neural network model to model the text sequence and extract semantic features; S32, Image data: Use convolutional neural networks to highlight important areas in the image for feature extraction; S33, audio data: use Mel frequency cepstral coefficients combined with convolutional neural network to extract audio features. The specific process is as follows: ① Audio preprocessing: Split the continuous audio signal into several short audio frames, each of which is 20-40 milliseconds long. There is a certain overlap between frames to ensure the continuity of the audio signal. Apply a window function to each audio frame for windowing processing, pre-emphasize the audio signal, and enhance the high-frequency part of the audio signal through a first-order high-pass filter to compensate for the attenuation of the speech signal in the high-frequency band and highlight the high-frequency resonance peak of the audio signal; ② Perform fast Fourier transform on the windowed audio frame to convert the time domain signal into the frequency domain signal to obtain the spectrum. The spectrum is converted from the linear frequency scale to the Mel frequency scale through the Mel filter bank. The Mel filter bank is a group of triangular bandpass filters uniformly distributed on the Mel frequency scale. The conversion relationship between the Mel frequency and the actual frequency is: Where f is the actual frequency. The logarithmic energy of each Mel filter output is taken to obtain the logarithmic Mel spectrum. The logarithmic Mel spectrum is subjected to discrete cosine transform to obtain the Mel frequency cepstrum coefficients. The first 13-20 coefficients are taken as the feature representation of the audio. ③ Use the calculated MFCC features as the input of the convolutional neural network to extract the audio features and convert them into feature vectors of fixed length.

5. The method of a multimodal cloud fusion service platform according to claim 1, characterized in that: In step S4, the process of evaluating and optimizing the fusion results is as follows: according to the specific application scenario, define specific evaluation indicators, then divide the data set into k equal parts, use k-1 parts for each training, and use the remaining part for verification, repeat k times, record the evaluation results of each fold and calculate the average value, count various evaluation indicators, analyze the performance of the model in different categories or different modes, identify the advantages and disadvantages of the model, analyze the confusion matrix, find the categories of model confusion and identify the direction that needs improvement.

6. The method of a multimodal cloud fusion service platform according to claim 1, characterized in that: In step S5, the platform is deployed to the cloud using containerized deployment technology, and a continuous integration and continuous deployment process is implemented to ensure that updates and improvements to the platform are released online simultaneously.

7. The method of a multimodal cloud fusion service platform according to claim 1, characterized in that: In step S6, the user uploads multimodal data and provides a visual data upload progress bar to let the user know the status and progress of the upload, and preview and verify the uploaded data in real time to ensure that the quality and format of the data meet the requirements; in the query service results, a collaborative filtering algorithm is used to provide the user with query results that meet their needs based on the user's historical queries and preferences.

Citation Information

Patent Citations

  • Emotion recognition method and system based on audio-visual feature correlation fusion

    CN111274955A

  • Multi-modal fusion-based subject portrait construction method, recommendation method and device

    CN118155031A

  • Contract question and answer method based on multiple modes

    CN118656469A