Custom AI interaction model-based gift personalized customization method

By combining multimodal data and custom AI interaction models to extract and process user emotional characteristics, the problem that the existing technology is difficult to accurately understand user emotions in various emotional complex situations is solved, and more accurate and personalized gift recommendations are achieved, improving user experience.

CN120219035AInactive Publication Date: 2025-06-27BEIJING BORUNYUAN TRADING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277937.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to accurately understand the user's true intentions and emotional needs in a variety of complex and intertwined emotions, resulting in a deviation between the gift recommendation results and the user's expected personalized needs.

Method used

By combining multimodal data (such as images, speech, and text), using a custom AI interaction model, text, image and speech emotional features are extracted and weighted averaged, and further processing is used to extract more abstract emotional features, and finally output user emotional categories through the Softmax classifier.

Benefits of technology

It realizes a comprehensive capture and understanding of the multi-dimensional emotional state of users, provides more accurate and personalized gift recommendations, and improves the accuracy and user experience of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219035A_ABST
    Figure CN120219035A_ABST
Patent Text Reader

Abstract

The invention relates to a gift personalized customization method based on a user-defined AI interaction model, in particular to the technical field of personalized electronic commerce, and aims to comprehensively capture and understand a multi-dimensional emotional state of a user through combination of multi-modal data (such as images, voices and texts) so as to provide more accurate and personalized recommendation. The multi-dimensional sentiment analysis breaks through the limitation of traditional single text analysis, subtle changes and diversity of user sentiments can be more deeply captured, the accuracy of a recommendation system and the user experience are improved, and the method is suitable for recommendation scenes needing high personalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of personalized e-commerce, and more specifically, to a gift personalization customization method based on a custom AI interaction model. Background Art

[0002] With the rapid development of e-commerce and online platforms, personalized recommendation systems play an increasingly important role in the user experience. Especially in the field of gift recommendation, how to make accurate recommendations according to the emotional needs of users has become a key factor in improving user satisfaction and purchase conversion rate. Existing sentiment analysis technologies attempt to identify the emotional state of users by analyzing multi-modal data such as users' text, voice, and images, so as to recommend suitable gifts for them. However, due to the limitations of sentiment analysis technologies, especially in situations where multiple emotions are intertwined, existing AI models are difficult to accurately understand the true intentions and emotional needs of users, resulting in a deviation between the recommended results and the personalized needs expected by users, and unable to truly achieve accurate matching.

[0003] Existing gift personalization customization methods based on custom AI interaction models, although achieving a combination of sentiment analysis and personalized recommendation to a certain extent, still face many challenges. First, due to the insufficient accuracy and precision of sentiment analysis models, especially in complex scenarios with multiple emotions intertwined, the models are unable to effectively distinguish the true emotional needs of users. For example, when a user expresses both happy and doubtful emotions, existing models may only judge their emotions as "happy" or "hesitant", ignoring the potential complex emotions of the user, thus resulting in the recommended gifts not meeting the deep needs of the user; second, existing sentiment analysis methods usually rely on fixed sentiment classification criteria, lacking dynamic capture of factors such as emotional intensity and emotional change, and unable to respond sensitively to the subtle changes in users' emotions, affecting the personalization and accuracy of gift recommendations. Summary of the Invention

[0004] In view of the technical problems existing in the prior art, the present invention provides a gift personalization customization method based on a custom AI interaction model, which combines multi-modal data (such as images, voice, and text), can comprehensively capture and understand the multi-dimensional emotional state of users, and thus provide more accurate and personalized recommendations to solve the problems proposed in the above background art.

[0005] The technical solution of the present invention to solve the above technical problems is as follows: Specifically, it includes the following steps: Step S1, receive the text data, image data, and voice data input by the user at the user interface terminal, and perform corresponding preprocessing on different data, including cleaning useless characters, removing stop words, and sub-word segmentation on the text data, and adding [CLS] and [SEP] tags to generate a standardized text sequence; perform scaling, cropping, and pixel normalization on the image data to generate a standardized image; perform MFCC audio feature extraction and normalization on the voice data input by the user to generate a standardized voice spectrum;

[0006] Step S2, input the standardized text sequence into the hidden layer of the BERT model, and extract the vector at the position of the last layer [CLS] as the text emotion feature vector; input the standardized image into the convolutional neural network CNN to extract the image emotion feature vector; input the standardized voice spectrum into the Wav2Vec model to extract the voice emotion feature vector;

[0007] Step S3, perform weighted averaging on the text, image, and voice emotion feature vectors to generate a multi-modal fusion feature, and further process it using the Transformer model to extract more abstract emotion features;

[0008] Step S4, input the emotion features extracted in S3 into the Softmax classifier to output the user emotion category, and the emotion category includes positive emotion, negative emotion, and neutral emotion.

[0009] In a preferred embodiment, in step S1, the text data includes the input comments, conversation content, and description content for personalized customization of gifts. Use the pre-trained BERT model to represent each word as a vector, thereby generating a standardized text sequence.

[0010] In a preferred embodiment, the image data includes the gift pictures and facial expression images uploaded by the user; use bilinear interpolation to adjust the image size, and the specific calculation formula is:

[0011] I resized (x,y)=∑ i,j I(i,j)·max(0,1-|x-i|)·max(0,1-|y-i|);

[0012] where, I(i,j) represents the pixel value of the original image, and (x,y) represents the coordinates of the scaled image;

[0013] Crop the target area from the scaled image, and the cropping formula is:

[0014] I cropped =I resized (x c:x c + ω,y c :y c + h);

[0015] Among them, (x c , y c ) represents the starting coordinates of cropping, ω and h respectively represent the target width and height;

[0016] Using the normalization formula, the pixel values of the image are converted from the original range to a specific range [0, 1], thereby generating a normalized image. The specific calculation formula is:

[0017]

[0018] Among them, I represents the pixel value of the original image.

[0019] In a preferred embodiment, after receiving the voice data, the specific operations for preprocessing the voice data are as follows:

[0020] A1. Process it using the pre-emphasis formula. The specific calculation formula is:

[0021] s pre (n) = s(n) - αs(n - 1), α ∈ [0.95, 0.97];

[0022] Among them, s pre (n) represents the nth sampling point of the pre-emphasized voice signal, s(n) represents the nth sampling point of the original voice signal, and α represents the pre-emphasis coefficient;

[0023] A2. Divide the pre-emphasized voice signal into short time segments. Define each frame as 25 ms and the frame shift as 10 ms, and perform windowing operations on it. The specific formula is:

[0024]

[0025] Among them, N represents the frame length, w(n) represents the window function value of the nth sampling point, and n represents the index of the current sampling point;

[0026] A3. Use the short-time Fourier transform formula to convert the time-domain signal into a frequency-domain signal. The specific calculation formula is:

[0027]

[0028] Among them, X(k) represents the complex spectrum value of the kth frequency component, s(n) represents the nth sampling point of the original voice signal, w(n) represents the window function value of the nth sampling point, k represents the index of the frequency component, and N represents the frame length;

[0029] A4. Convert the linear frequency to Mel frequency using the Mel filter bank. The specific calculation formula is:

[0030]

[0031] where mel(f) represents the value obtained by converting the linear frequency f to Mel frequency, and f represents the linear frequency;

[0032] A5. Extract MFCC coefficients using the discrete cosine transform formula. The specific calculation formula is:

[0033]

[0034] where c m represents the m-th MFCC coefficient, (E k ) represents the energy of the k-th Mel filter, M represents the number of Mel filters, and m represents the index of the MFCC coefficient;

[0035] A6. Perform Z-score normalization on the MFCC features to generate a normalized speech spectrum. The specific calculation formula is:

[0036]

[0037] where represents the m-th MFCC coefficient after normalization, c m represents the original m-th MFCC coefficient, μ m represents the mean of the m-th MFCC coefficient in the training set, and σ m represents the standard deviation of the m-th MFCC coefficient in the training set.

[0038] In a preferred embodiment, in step S3, the text, image, and speech emotion feature vectors are all mapped to the same dimension through a fully connected layer. Secondly, the weight of the text emotion feature vector w1 = 0.4, the weight of the image emotion feature vector w2 = 0.3, and the weight of the speech emotion feature vector w3 = 0.3 are set.

[0039] In a preferred embodiment, the calculation formula for the weighted average is:

[0040] h fused = w1·h text + w2·h image + w3·h audio ;

[0041] where h text , h image , h audiorespectively represent the text emotion modality feature vector, the image emotion modality feature vector, and the speech emotion modality feature vector, and w1, w2, and w3 respectively represent the fusion weights of each modality, h fused represents the fused comprehensive feature vector.

[0042] In a preferred embodiment, in the step S3, the specific processing steps for extracting more abstract emotion features are as follows:

[0043] A1. Directly input the fused comprehensive feature vector h fused , and add positional encoding;

[0044] A2. Use the self-attention mechanism to capture the internal relationships of the features, and the calculation formula is:

[0045]

[0046] where Q, K, and V respectively represent the query, key, and value matrices, which are obtained by input projection, and d k represents the dimension of the key vector;

[0047] A3. Take the feature at the position of the [CLS] token as the output emotion feature.

[0048] In a preferred embodiment, in the step S4, the specific steps for emotion classification by the Softmax classifier are as follows:

[0049] A1. Use the linear transformation formula to map the features to the category space to obtain the unnormalized category scores, and the calculation formula is:

[0050] z = h emotion ·W + b;

[0051] where z represents the unnormalized category scores after linear transformation, h emotion represents the input emotion feature vector, W represents the weight matrix, and b represents the bias term vector;

[0052] A2. Use the Softmax function to normalize z to obtain the probability distribution The calculation formula is:

[0053]

[0054] where z i represents the unnormalized score of the i-th category, exp(z i ) represents taking the exponential of z i to convert it to a positive number, represents the sum of the exponential values of all categories, represents the probability value of the i-th category, and C represents the total number of emotion categories;

[0055] A3. Obtain the probability distribution Take the category corresponding to the maximum probability as the final prediction result, and the calculation formula is:

[0056]

[0057] Among them, pred_class represents the index of the predicted sentiment category, argmax represents the index of the category with the maximum probability, and at the same time, according to the index, look up the preset sentiment category label table to output the user sentiment category.

[0058] In a preferred embodiment, the sentiment category label table is specifically set as follows: 1 represents "happy", 1 represents "angry", 2 represents "satisfied", 3 represents "excited", 4 represents "angry", 5 represents "sad", 6 represents "frustrated", 7 represents "calm", 8 represents "indifferent".

[0059] In a preferred embodiment, the positive emotions include: happy, satisfied, excited; the negative emotions include: angry, angry, sad, frustrated; the neutral emotions include: calm, indifferent.

[0060] The beneficial effects of the present invention are: By combining multi-modal data (such as images, voices, texts), the present invention can comprehensively capture and understand the multi-dimensional emotional states of users, so as to provide more accurate and personalized recommendations. This multi-dimensional emotional analysis breaks through the limitations of traditional single-text analysis, can more deeply capture the subtle changes and diversity of users' emotions, improves the accuracy of the recommendation system and the user experience, and is applicable to highly personalized recommendation scenarios. Brief Description of the Drawings

[0061] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0063] In the description of the present application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.

[0064] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily to be construed as more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not described in detail so as not to obscure the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed in the present application.

[0065] Embodiment 1

[0066] This embodiment provides a method for personalized customization of gifts based on a custom AI interaction model as shown in Figure 1 The method specifically includes the following steps: Step S1: Receive the text data, image data, and voice data input by the user at the user interface terminal, and perform corresponding preprocessing on different data, including cleaning useless characters (such as punctuation marks, special symbols, etc.), removing stop words (such as "de", "le", "shi", etc.) from the text data, and performing sub-word segmentation and adding [CLS] and [SEP] tags (splitting the text into words or phrases) to generate a standardized text sequence; performing scaling, cropping, and pixel normalization processing on the image data to generate a standardized image; performing MFCC audio feature extraction and normalization processing on the voice data input by the user to generate a standardized voice spectrum;

[0067] Step S2: Input the standardized text sequence into the hidden layer of the BERT model, and extract the vector at the position of the last layer [CLS] as the text sentiment feature vector; input the standardized image into the convolutional neural network CNN to extract the image sentiment feature vector; input the standardized voice spectrum into the Wav2Vec model to extract the voice sentiment feature vector;

[0068] Step S3: Perform weighted averaging on the text, image, and voice sentiment feature vectors to generate a multi-modal fusion feature, and further process it using the Transformer model to extract more abstract sentiment features;

[0069] Step S4: Input the sentiment features extracted in S3 into the Softmax classifier to output the user sentiment category, where the sentiment category includes positive sentiment, negative sentiment, and neutral sentiment.

[0070] In this embodiment, it should be specifically noted that in step S1, the text data includes the input comments, conversation content, and description content of the personalized customization of gifts. The comments refer to the evaluation content of users on goods or services; the conversation content refers to the customer service chat records or the conversation between users and the system; the description content of the personalized customization of gifts refers to the specific requirements of users for customized gifts. Through the above text data, valuable information can be provided for subsequent analysis of their preferences and personalized needs, so as to provide more accurate recommendation or customization services; in order to convert the text into a digital form that can be understood by a computer, this application uses a pre-trained BERT model to represent each word as a vector, thereby generating a standardized text sequence. This operation can capture the relationships and semantic information between words.

[0071] The image data includes the gift pictures uploaded by users and facial expression images; use bilinear interpolation to adjust the image size, and the specific calculation formula is:

[0072] I resized (x,y) = ∑ i,j I(i,j)·max(0,1 - |x - i|)·max(0,1 - |y - i|);

[0073] where I(i,j) represents the pixel value of the original image, and (x,y) represents the coordinates of the scaled image;

[0074] Crop the target area from the scaled image, and the target area usually refers to the central area. The cropping formula is:

[0075] I cropped = I resized (x c :x c + ω,y c :y c + h);

[0076] where (x c ,y c ) represents the starting coordinates of the cropping, and ω and h respectively represent the target width and height;

[0077] Use the normalization formula to convert the pixel values of the image from the original range (such as [0,255]) to a specific range [0,1] to accelerate model convergence and improve performance, thereby generating a standardized image. The specific calculation formula is:

[0078]

[0079] Among them, I represents the pixel value of the original image;

[0080] After receiving the voice data, the specific operations for preprocessing the voice data are as follows:

[0081] A1. Process it using the pre-emphasis formula to enhance high-frequency signals, balance the spectrum, reduce the attenuation of high-frequency signals, and make the signal more suitable for subsequent processing. The specific calculation formula is:

[0082] s pre (n) = s(n) - αs(n - 1), α ∈ [0.95, 0.97];

[0083] Among them, s pre (n) represents the nth sampling point of the pre-emphasized voice signal, s(n) represents the nth sampling point of the original voice signal, and α represents the pre-emphasis coefficient, usually taking a value from 0.95 to 0.97, which is used to enhance high-frequency signals;

[0084] A2. Divide the pre-emphasized voice signal into short time periods (frames), define each frame as 25 ms (when the sampling rate is 16 kHz, the frame length is 400 sampling points), and the frame shift is 10 ms (when the sampling rate is 16 kHz, the frame shift is 160 sampling points), and perform a windowing operation on it. The purpose is to make the signal more continuous within the frame by smoothing the frame edges. The specific formula is:

[0085]

[0086] Among them, N represents the frame length (such as 400 sampling points), w(n) represents the window function value of the nth sampling point, and n represents the index of the current sampling point;

[0087] A3. Use the short-time Fourier transform formula to convert the time-domain signal into a frequency-domain signal and analyze the spectral characteristics of each frame. The specific calculation formula is:

[0088]

[0089] Among them, X(k) represents the complex spectral value of the kth frequency component, s(n) represents the nth sampling point of the original voice signal, w(n) represents the window function value of the nth sampling point, k represents the index of the frequency component, and N represents the frame length (i.e., the number of sampling points per frame);

[0090] A4. Use the Mel filter bank to convert the linear frequency to the Mel frequency. The specific calculation formula is:

[0091]

[0092] Among them, mel(f) represents the value obtained by converting the linear frequency f to the mel frequency, and f represents the linear frequency (unit: Hz);

[0093] A5. Extract MFCC coefficients (take the first 13 dimensions) using the discrete cosine transform formula. The specific calculation formula is:

[0094]

[0095] Among them, c m represents the m-th MFCC coefficient, (E k ) represents the energy of the k-th mel filter, M represents the number of mel filters (usually 40), and m represents the index of the MFCC coefficient (take the first 13 dimensions);

[0096] A6. Perform Z-score normalization on the MFCC features to generate a normalized speech spectrum. The specific calculation formula is:

[0097]

[0098] Among them, represents the m-th MFCC coefficient after normalization, c m represents the original m-th MFCC coefficient, μ m represents the mean of the m-th MFCC coefficient in the training set, and σ m represents the standard deviation of the m-th MFCC coefficient in the training set.

[0099] In this embodiment, it should be specifically noted that in step S3, the text, image, and speech emotion feature vectors are all mapped to the same dimension through a fully connected layer. For example, if the text feature is 768 dimensions, the image feature is 2048 dimensions, and the speech feature is 1024 dimensions, they can be mapped to 512 dimensions respectively. Secondly, set the weight w1 of the text emotion feature vector to 0.4, the weight w2 of the image emotion feature vector to 0.3, and the weight w3 of the speech emotion feature vector to 0.3;

[0100] The calculation formula for weighted average is:

[0101] h fused = w1·h text + w2·h image + w3·h audio ;

[0102] Among them, h text , h image , h audio represent the text emotion modality feature vector, the image emotion modality feature vector, and the speech emotion modality feature vector respectively. w1, w2, and w3 represent the fusion weights of each modality, and h fusedRepresents the fused comprehensive feature vector, with the same dimension as the input feature;

[0103] In step S3, the specific processing steps for extracting more abstract emotional features are as follows:

[0104] A1. Directly input the fused comprehensive feature vector h fused (with the shape of [1, d]), and add positional encoding;

[0105] A2. Use the self-attention mechanism to capture the internal relationships of the features, and the calculation formula is:

[0106]

[0107] Among them, Q, K, and V respectively represent the query, key, and value matrices, which are obtained by projecting the input, and d k represents the dimension of the key vector;

[0108] A3. Take the feature at the position of the [CLS] token as the output emotional feature.

[0109] In this embodiment, specifically, it should be noted that in step S4, the specific steps of the emotional classification by the Softmax classifier are as follows:

[0110] A1. Use the linear transformation formula to map the features to the category space to obtain the unnormalized category scores, and the calculation formula is:

[0111] z = h emotion ·W + b (with the shape of [1 * C]);

[0112] Among them, z represents the unnormalized category scores after linear transformation, represents a vector of length C, and each element z i represents the score corresponding to the category, h emotion represents the input emotional feature vector, W represents the weight matrix that maps the feature vector to the category space, and b represents the bias term vector;

[0113] A2. Use the Softmax function to normalize z to obtain the probability distribution The calculation formula is:

[0114]

[0115] Among them, z i represents the unnormalized score of the i-th category, exp(z i ) represents taking the exponential of z i to convert it to a positive number, represents the sum of the exponential values of all categories, which is used for normalization, represents the probability value of the i-th category, and C represents the total number of emotional categories;

[0116] A3. Obtain the probability distribution Take the category corresponding to the maximum probability in the probability distribution as the final prediction result. The calculation formula is as follows:

[0117]

[0118] Among them, pred_class represents the index of the predicted sentiment category, argmax represents the index that returns the category with the maximum probability. At the same time, according to the index, look up the preset sentiment category label table to output a more accurate analysis result of the user's sentiment category, so as to achieve a more comprehensive understanding of the user's sentiment state, so that the AI model can more accurately perceive the user's sentiment;

[0119] The specific settings of the sentiment category label table are as follows: 1 represents "happy", 1 represents "angry", 2 represents "satisfied", 3 represents "excited", 4 represents "angry", 5 represents "sad", 6 represents "frustrated", 7 represents "calm", 8 represents "indifferent";

[0120] Positive emotions include: happy, satisfied, excited; negative emotions include: angry, angry, sad, frustrated; neutral emotions include: calm, indifferent.

[0121] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0122] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks

[0124] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0125] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0126] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0127] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A gift customization method based on a custom AI interaction model, characterized in that: The specific steps include: Step S1, receiving text data, image data and voice data input by the user in the user interface terminal, and performing corresponding preprocessing on different data, including cleaning useless characters, removing stop words, and subword segmentation of text data and adding [CLS] and [SEP] tags to generate a standardized text sequence; scaling, cropping and pixel normalization of image data to generate a standardized image; Perform MFCC audio feature extraction and standardization on the voice data entered by the user to generate a standardized voice spectrum; Step S2: Input the standardized text sequence into the hidden layer of the BERT model, and extract the vector at the position of [CLS] in the last layer as the text sentiment feature vector; input the standardized image into the convolutional neural network CNN to extract the image sentiment feature vector; input the standardized speech spectrum into the Wav2Vec model to extract the speech sentiment feature vector; Step S3: weighted average the text, image and speech emotion feature vectors to generate multimodal fusion features, which are further processed using the Transformer model to extract more abstract emotion features; Step S4: input the emotion features extracted in S3 into the Softmax classifier, and output the user emotion category, which includes positive emotion, negative emotion and neutral emotion.

2. According to claim 1, a gift customization method based on a custom AI interaction model is characterized in that: In step S1, the text data includes input comments, conversation content, and description content of personalized gift customization. The pre-trained BERT model is used to represent each word as a vector, thereby generating a standardized text sequence.

3. A gift customization method based on a custom AI interaction model according to claim 2, characterized in that: The image data includes gift pictures and facial expression images uploaded by users; the image size is adjusted using bilinear interpolation, and the specific calculation formula is: I resized (x,y)=∑ i,j I(i,j)·max(0,1-|x-i|)·max(0,1-|y-i|); Among them, I(i,j) represents the pixel value of the original image, and (x,y) represents the coordinates of the scaled image; Cut out the target area from the scaled image, and the cropping formula is: I cropped =I resized (x c :x c +ω,y c :y c +h); Among them, (x c ,y c ) represents the starting coordinates of the crop, ω and h represent the target width and height respectively; Using the normalization formula, the pixel values ​​of the image are converted from the original range to a specific range [0,1] to generate a standardized image. The specific calculation formula is: Where I represents the pixel value of the original image.

4. According to claim 3, a gift personalization customization method based on a custom AI interaction model is characterized in that: After receiving the voice data, the specific operations for preprocessing the voice data are as follows: A1. Use the pre-emphasis formula to process it. The specific calculation formula is: s pre (n)=s(n)-αs(n-1),α∈[0.95,0.97]; Among them, s pre (n) represents the nth sampling point of the pre-emphasized speech signal, s(n) represents the nth sampling point of the original speech signal, and α represents the pre-emphasis coefficient; A2. Divide the pre-emphasized speech signal into short time periods, define each frame as 25ms, frame shift as 10ms, and perform windowing operation on it. The specific formula is: Where N represents the frame length, w(n) represents the window function value of the nth sampling point, and n represents the index of the current sampling point; A3. Use the short-time Fourier transform formula to convert the time domain signal into the frequency domain signal. The specific calculation formula is: Where X(k) represents the complex spectrum value of the kth frequency component, s(n) represents the nth sampling point of the original speech signal, w(n) represents the window function value of the nth sampling point, k represents the index of the frequency component, and N represents the frame length; A4. Use the Mel filter bank to convert the linear frequency into Mel frequency. The specific calculation formula is: Wherein, mel(f) represents the value of converting the linear frequency f into the Mel frequency, and f represents the linear frequency; A5. Use the discrete cosine transform formula to extract the MFCC coefficients. The specific calculation formula is: Among them, c m represents the mth MFCC coefficient, (E k ) represents the energy of the kth Mel filter, M represents the number of Mel filters, and m represents the index of the MFCC coefficient; A6. Perform Z-score normalization on the MFCC features to generate a standardized speech spectrum. The specific calculation formula is: in, represents the mth MFCC coefficient after normalization, c m Represents the original m-th MFCC coefficient, μ m represents the mean of the mth MFCC coefficient in the training set, σ m Represents the standard deviation of the mth MFCC coefficient in the training set.

5. According to claim 4, a gift personalization customization method based on a custom AI interaction model is characterized in that: In the step S3, the text, image and speech emotion feature vectors are all mapped to the same dimension through a fully connected layer. Next, the weight w1 of the text emotion feature vector is set to 0.4, the weight w2 of the image emotion feature vector is set to 0.3, and the weight w3 of the speech emotion feature vector is set to 0.

3.

6. A gift customization method based on a custom AI interaction model according to claim 5, characterized in that: The calculation formula of the weighted average is: h fused =w1·h text +w2·h image +w3·h audio ; Among them, h text 、h image 、h audio They represent the text emotion modality feature vector, image emotion modality feature vector and speech emotion modality feature vector respectively, w1, w2, w3 represent the fusion weights of each modality respectively, and h fused Represents the integrated feature vector after fusion.

7. A gift customization method based on a custom AI interaction model according to claim 6, characterized in that: In step S3, the specific processing steps for extracting more abstract emotional features are as follows: A1. Directly input the fused comprehensive feature vector h fused , and add position coding; A2. Use the self-attention mechanism to capture the internal relationship of features. The calculation formula is: Among them, Q, K, and V represent query, key, and value matrices respectively, which are obtained by input projection. k represents the dimension of the key vector; A3. Take the features of the [CLS] tag position as the output sentiment features.

8. A gift customization method based on a custom AI interaction model according to claim 7, characterized in that: In step S4, the specific steps of sentiment classification of the Softmax classifier are as follows: A1. Use the linear transformation formula to map the features to the category space and obtain the unnormalized category score. The calculation formula is: z=h emotion ·W+b; Among them, z represents the unnormalized category score after linear transformation, h emotion represents the input sentiment feature vector, W represents the weight matrix, and b represents the bias term vector; A2. Use the Softmax function to normalize z and get the probability distribution The calculation formula is: Among them, z i represents the unnormalized score of the ith category, exp(z i ) indicates the i Take the exponent, convert it to a positive number, It means summing the index values ​​of all categories. represents the probability value of the i-th category, and C represents the total number of emotion categories; A3. Get probability distribution The category corresponding to the maximum probability is taken as the final prediction result, and the calculation formula is: Among them, pred class represents the predicted sentiment category index, argmax represents the index of the category with the highest probability of returning, and at the same time, the preset sentiment category label table is searched according to the index to output the user sentiment category.

9. A gift customization method based on a custom AI interaction model according to claim 8, characterized in that: The specific settings of the emotion category label table are as follows: 1 means "happy", 1 means "angry", 2 means "satisfied", 3 means "excited", 4 means "angry", 5 means "sad", 6 means "depressed", 7 means "calm", and 8 means "indifferent".

10. A gift customization method based on a custom AI interaction model according to claim 9, characterized in that: Positive emotions include: happiness, satisfaction, excitement; negative emotions include: anger, rage, sadness, frustration; neutral emotions include: calmness and indifference.