A cross-modal semantic parsing method
By performing visual and auditory modal analysis on the source video and fusing it with textual information, the problem of low semantic parsing accuracy caused by intermodal differences is solved, achieving higher accuracy and adaptability in cross-modal semantic parsing.
Patent Information
- Application Number
- CN202510908081.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Due to the significant differences in information structure and expression methods among different modalities, existing cross-modal semantic parsing methods perform poorly when dealing with modal conflicts, semantic ambiguities, or missing modal information, resulting in low semantic parsing accuracy.
The source video is analyzed by invoking visual and auditory modal parsing strategies to obtain visual and auditory parsing information. The results are then compared and analyzed in stages, the modal weights are dynamically adjusted, and the text parsing information is fused to obtain semantic parsing results.
It improves the accuracy and adaptability of cross-modal semantic parsing, enabling a more accurate understanding of sentiment expression in multimodal data.
Smart Images

Figure CN120408537B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a cross-modal semantic analysis method. BACKGROUND
[0002] In the customer service interaction process, the intention of the user is often expressed through multiple modalities, including text input, voice dialogue, images (such as screenshots, photos), etc. Different modalities (such as text, voice, and images) have great differences in expression and semantic content. Due to the differences in information structure and expression between modalities, it is difficult to establish the semantic correspondence between different modalities in cross-modal semantic analysis, thereby causing the alignment problem of cross-modal data. In addition, due to the lack of dynamic perception ability of the consistency of emotional expression and the difference of semantic expression between modalities, the existing cross-modal semantic analysis performs poorly in dealing with modality conflicts, semantic ambiguities or missing modal information, ultimately resulting in low semantic analysis accuracy.
[0003] In summary, the prior art has the technical problem of low semantic analysis accuracy due to the great differences in information structure and expression between modalities. SUMMARY
[0004] The purpose of the present application is to provide a cross-modal semantic analysis method to solve the technical problem of low semantic analysis accuracy due to the great differences in information structure and expression between modalities in the prior art.
[0005] In view of the above problems, the present application provides a cross-modal semantic analysis method, wherein the cross-modal semantic analysis method comprises: calling a visual modality analysis strategy to analyze a source video to obtain visual analysis information; calling an auditory modality analysis strategy to analyze source audio extracted from the source video to obtain auditory analysis information; performing stage-by-stage comparative analysis on the visual analysis information and the auditory analysis information to obtain a sentiment comparison result; if the sentiment comparison result reaches a predetermined comparison constraint, analyzing a modality analysis database to obtain real-time modality weight distribution; obtaining text analysis information of source text, and combining the real-time modality weight distribution to fuse the text analysis information, the visual analysis information and the auditory analysis information to obtain a semantic analysis result.
[0006] Optionally, the source video is subjected to dynamic image compression processing to obtain a compressed video; the compressed video is subjected to detection analysis according to a shot detection mechanism in the visual modality analysis strategy to obtain a detection result; a first image sequence corresponding to a first shot in the detection result is extracted, and the first image sequence is subjected to enhancement fusion processing to obtain a first target image; the first target image is subjected to emotion analysis according to an emotion analysis mechanism in the visual modality analysis strategy to obtain first emotion information; a first video segment corresponding to the first shot in the source video is matched, and the visual analysis information is established based on a first mapping relationship between the first video segment and the first emotion information.
[0007] Optionally, adjacent image groups of the compressed video are obtained; adjacent displacement values of a first image and a second image in the adjacent image groups are calculated according to the shot detection mechanism; if the adjacent displacement values do not conform to a predetermined displacement limit value, the first image and the second image are different shots, and the detection result is formed.
[0008] Optionally, a first feature parameter set of the first target image is collected in multiple dimensions; a first emotion dimension in a predetermined emotion dimension is extracted, wherein the predetermined emotion dimension is stored in the emotion analysis mechanism; a first dimension coefficient of the first emotion dimension is obtained in combination with the first feature parameter set; and the first emotion information is established based on the first dimension coefficient.
[0009] Optionally, a sound signal of the source audio is obtained, and a sound intensity time sequence is analyzed; the sound intensity time sequence is subjected to random segmentation to obtain a segmentation result; a first time domain feature parameter of a first time sequence in the segmentation result is collected; a first frequency domain feature parameter of a first frequency spectrum obtained by fast Fourier transform processing of the first time sequence is collected; a sound emotion prediction model in the auditory modality analysis strategy is activated to perform prediction analysis on the first time domain feature parameter and the first frequency domain feature parameter to obtain first prediction emotion information; a first audio segment corresponding to the first time sequence in the source audio is matched, and the auditory analysis information is established based on a second mapping relationship between the first audio segment and the first prediction emotion information.
[0010] Optionally, the sound emotion prediction model includes three prediction channels of the predetermined emotion dimension, and the predetermined emotion dimension includes emotion valence, arousal, and dominance.
[0011] Optionally, a sound emotion dictionary is constructed, and a first training data set of a first emotion dimension in the predetermined emotion dimension is established based on the sound emotion dictionary; a first prediction channel is obtained by supervised learning on the first training data set, and the sound emotion prediction model is composed; wherein the first training data set includes first training time domain feature parameters, first training frequency domain feature parameters and first training emotion information.
[0012] Optionally, a modal analysis sample is obtained by sampling in the modal analysis database; a review result is obtained by manual emotion review on the modal analysis sample; a modal analysis accuracy rate is obtained by analyzing the review result, wherein the modal analysis accuracy rate includes a visual analysis accuracy rate and an auditory analysis accuracy rate; the real-time modal weight distribution is obtained by variation weighting analysis on the visual analysis accuracy rate and the auditory analysis accuracy rate.
[0013] Optionally, a processing result is obtained by cleaning and word segmentation processing on the source text; a feature vector is obtained by feature analysis on the processing result based on a bag-of-words model principle; the text analysis information is obtained by emotion analysis on the feature vector in combination with a vocabulary emotion dictionary.
[0014] Optionally, any vocabulary in the feature vector is extracted; any score of the any vocabulary is matched and obtained in the vocabulary emotion dictionary, wherein the any score includes three scores of the predetermined emotion dimension; any frequency of the any vocabulary is obtained according to the feature vector, and any emotion score of the any vocabulary is obtained in combination with the three scores; a mean value of the any emotion score is taken to compose the text analysis information.
[0015] The technical solutions provided in the application have at least the following beneficial effects:
[0016] The source video is parsed by calling a visual modal analysis strategy to obtain visual analysis information; the source audio extracted from the source video is parsed by calling an auditory modal analysis strategy to obtain auditory analysis information; the visual analysis information and the auditory analysis information are subjected to stage-by-stage comparative analysis to obtain a sentiment comparison result; if the sentiment comparison result reaches a predetermined comparison constraint, real-time modal weight distribution is obtained by analyzing a modal analysis database; text analysis information of source text is obtained, and the text analysis information, the visual analysis information and the auditory analysis information are fused in combination with the real-time modal weight distribution to obtain a semantic analysis result. That is, visual content in the source video is parsed to extract visual analysis information; source audio in the source video is parsed to obtain auditory analysis information; the visual analysis information and the auditory analysis information are subjected to stage-by-stage comparative analysis to obtain a sentiment comparison result; by dynamically adjusting the weight of each modal according to the sentiment comparison result, the analysis information of text, vision and hearing is fused according to the real-time modal weight distribution to obtain a semantic analysis result, thereby improving the accuracy and adaptability of cross-modal semantic analysis.
[0017] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only exemplary, and other drawings can be obtained by those skilled in the art without creating any creative labor on the basis of the provided drawings.
[0019] Figure 1 A flowchart of a cross-modal semantic analysis method of the present application.
[0020] Figure 2 A flowchart of obtaining visual analysis information in a cross-modal semantic analysis method of the present application. DETAILED DESCRIPTION
[0021] The application provides a cross-modal semantic parsing method, which solves the technical problem of low semantic parsing accuracy caused by large differences in information structure and expression methods between different modalities in the prior art. By parsing the visual content in the source video, visual parsing information is extracted; the source audio in the source video is parsed to obtain auditory parsing information; the visual parsing information and the auditory parsing information are compared and analyzed in stages to obtain a sentiment comparison result; by analyzing the sentiment comparison result, the weight of each modality is dynamically adjusted, and the parsing information of the text, the vision and the hearing is fused according to the real-time modality weight distribution to obtain a semantic parsing result, thereby improving the accuracy and adaptability of cross-modal semantic parsing.
[0022] The technical solutions in the application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the application, rather than all the embodiments of the application. It should be understood that the application is not limited by the example embodiments described herein. Based on the embodiments of the application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the application. In addition, it should be noted that, for the convenience of description, only parts related to the application are shown in the drawings, not all.
[0023] Embodiments, please refer to the drawings Figure 1 The application provides a cross-modal semantic parsing method, which specifically includes the following steps:
[0024] S100: retrieving a visual modality parsing strategy to parse the source video to obtain visual parsing information.
[0025] Further, as shown in the accompanying Figure 2 The application S100 includes:
[0026] S110: performing dynamic image compression processing on the source video to obtain a compressed video; S120: detecting and analyzing the compressed video according to a shot detection mechanism in the visual modality parsing strategy to obtain a detection result; S130: extracting a first image sequence corresponding to a first shot in the detection result, and performing enhancement fusion processing on the first image sequence to obtain a first target image; S140: performing sentiment analysis on the first target image according to a sentiment analysis mechanism in the visual modality parsing strategy to obtain first sentiment information; S150: matching a first video segment corresponding to the first shot in the source video, and establishing the visual parsing information based on a first mapping relationship between the first video segment and the first sentiment information.
[0027] Specifically, the source video refers to the original video file, which usually contains various audio-visual information such as image frames, audio signals, and subtitles. Dynamic image compression processing is performed on the source video to reduce the data size of the video and improve the efficiency of subsequent processing. Through compression algorithms such as H.264, HEVC, etc., efficient compression is achieved by reducing redundant data in the video. The specific process is as follows: load the source video and apply dynamic image compression algorithms for compression processing. During compression, mechanisms such as inter-frame prediction, motion estimation, and entropy coding are used to effectively eliminate redundant information between adjacent frames in the video, thereby significantly reducing the data volume. For example, a voice and video conversation collected by a customer service center has a format of MP4, a resolution of 1920x1080 (full HD), a frame rate of 30fps, and a length of 3 minutes, with an original size of 900MB. Using the HEVC (H.265) compression algorithm, the compression parameter settings are as follows: the code rate control is set to constant quality (CRF=23); the encoding preset is set to slow. The compressed output video size is 180MB, with a compression ratio of about 80%, SSIM (Structural Similarity) = 0.96, and PSNR (Peak Signal-to-Noise Ratio) = 40.2dB, indicating that the image quality is basically preserved after compression.
[0028] According to the lens detection mechanism in the visual modal analysis strategy, the compressed video is analyzed to automatically identify different lenses in the video. The lens detection mechanism is used to identify the scene change points in the video, i.e., to determine whether the lens in the video has switched. Lens switching usually means changes in the scene, characters, perspective, or actions. According to the lens detection result, the first image sequence corresponding to the first lens is extracted, representing a scene in the video. The first image sequence is a series of consecutive image frames corresponding to the first image, each image frame is a static image of the video, and the image sequence is a sequence of these images arranged in chronological order, reflecting the dynamic content in the video.
[0029] The extracted first image sequence is subjected to enhancement fusion processing to improve the quality of the images, including denoising, enhancing contrast, and increasing image clarity. Through enhancement fusion processing, the details in the image sequence will be clearer, especially suitable for subsequent tasks such as emotion analysis and object recognition. Enhancement fusion processing improves the quality of images (such as enhancing contrast, denoising, and enhancing details) to generate clearer and more detailed target images. After enhancement fusion processing, the first target image is obtained, which is an optimized image usually with higher clarity and more details. For example, if the source video is shot in a low-light environment, denoising and brightness enhancement are performed on the image to make the customer service personnel's facial expressions more clearly visible.
[0030] The first target image is subjected to emotion analysis by an emotion analysis mechanism in a visual modal analysis strategy, to obtain first emotion information including emotion valence, arousal and dominance. A first video segment corresponding to the first shot is matched in the source video, and the first emotion information is matched with the corresponding first video segment in the source video to construct a first mapping relationship between the video and the emotion information. The mapping relationship describes the relationship between the video content and its emotional characteristics. The above steps are repeated for other shots in the detection result, and finally the mapping relationship between the video segment and the emotion information corresponding to each shot is obtained, which collectively constitutes the visual analysis information. The visual analysis information is the result obtained based on image or video analysis, and usually includes feature information, emotion information and their relationship.
[0031] Through dynamic image compression processing, the size of the video file is reduced, the processing efficiency is improved, and the demand for storage space and transmission bandwidth is reduced. Through the shot detection mechanism, the video content is automatically analyzed and organized, and the efficiency of video processing is improved. Through the emotion analysis mechanism, the emotional level of the video content is understood, which helps to analyze the video content more deeply.
[0032] Further, the present application further comprises the following steps:
[0033] S121: obtaining adjacent image groups of the compressed video; S122: calculating the adjacent displacement value of the first image and the second image in the adjacent image groups according to the shot detection mechanism; S123: if the adjacent displacement value does not meet the predetermined displacement limit value, the first image and the second image are different shots, and the detection result is formed.
[0034] Specifically, in the compressed video, the frames of the video are usually divided into multiple image groups, and adjacent image groups refer to two groups of image frames closely connected in time sequence. Adjacent image groups refer to two or more consecutive frames of images in a video, which are in adjacent time periods and reflect continuous scenes or actions in the video. That is, in the video processing process, two or more consecutive frames of images in the video are organized into adjacent image groups, and these frames belong to consecutive time points. The shot detection mechanism is a technology for analyzing the changes of shots (i.e. scenes) in a video, which identifies the shot transitions in the video by analyzing the changes between adjacent frames, and the shot transition is usually a sign of scene change.
[0035] The lens detection mechanism is used to estimate the motion between two adjacent frames of images. By comparing the differences between adjacent frames or images, it is determined whether the lens has shifted. The displacement between images is calculated by comparing the pixel changes between two frames, obtaining the adjacent displacement value of the first image and the second image. Through calculation, the degree of change between the first image and the second image is obtained, which is quantified as the adjacent displacement value, reflecting whether the lens moves, whether the picture content changes (such as cutting to another scene), etc.
[0036] The adjacent displacement value is measured by calculating the motion of image content, which can be completed by motion estimation techniques such as optical flow. The optical flow method estimates the motion information in the image based on the changes of image pixel points. The two consecutive frames of images (such as the first image and the second image) are preprocessed to extract key feature points such as corner points and edges. Using the optical flow method, the displacement vector of each pixel point is calculated by analyzing the brightness changes of the pixel points in the adjacent frames, including the displacement values in the horizontal and vertical directions. By integrating the optical flow vectors, the overall displacement value between adjacent images is calculated, and the average or maximum value of the displacement of all pixel points is calculated. If the adjacent displacement value exceeds the preset threshold value, it indicates that the two frames of images may belong to different lenses. If there is a lens switching in the recording (such as accidentally switching to other cameras), the adjacent displacement value will increase significantly. By detecting this abnormal displacement, these points are automatically marked and the lens switching points in the recording are quickly located. For example, consecutive frames are extracted from the video for calculation, such as the first frame and the second frame, the second frame and the third frame, etc. to form adjacent image groups. The customer service personnel in the first frame and the second frame are sitting steadily, and the background does not change, with a displacement value of 5 pixels; in the second frame and the third frame, the lens switches to the customer side, resulting in different backgrounds and expressions, with a displacement value of 40 pixels.
[0037] The calculated displacement value is compared with the preset displacement limit value (such as 20 pixels). If the displacement value is greater than the predetermined limit value, it indicates that the scene has changed and the lens has switched; if the displacement value is smaller, it indicates that the picture has not changed significantly and still belongs to the same lens. According to the preset displacement limit value, it is determined whether it is a lens switching, and the detection result marks the switching time in the source video. The detection result shows which frames in the video belong to the same lens and which frames belong to different lenses. Through the calculation of the adjacent displacement value, the lens switching points in the video are identified, and each lens is accurately located, thereby improving the quality of video analysis.
[0038] Further, the present application further comprises the following steps:
[0039] S141: multi-dimensionally collect a first feature parameter set of the first target image; S142: extract a first emotional dimension in a predetermined emotional dimension, wherein the predetermined emotional dimension is stored in the emotional analysis mechanism; S143: combine the first feature parameter set to obtain a first dimension coefficient of the first emotional dimension; and S144: construct the first emotional information based on the first dimension coefficient.
[0040] Specifically, the first target image is multi-dimensionally collected to obtain a first feature parameter set, including facial expressions, action units (such as raised eyebrows, raised corners of the mouth), color distribution, texture, etc. Multi-dimensional collection refers to extracting multiple different types of features from the first target image to comprehensively capture the emotional and state information of the image or scene. For example, the specific expressions of the face, such as raised corners of the mouth and raised eyebrows, are extracted through facial action units; the positions of the facial key points are detected through a deep learning model (such as VGGFace, ResNet) to determine the degree of mouth opening and eye closure.
[0041] From the collected multi-dimensional features, a first feature parameter set is selected and sorted, including various features related to emotional analysis, such as facial muscle activity, eye state, expression change, etc. For example: facial expression: raised eyebrows (AU1) = 0.8, raised corners of the mouth (AU12) = 0.6; hue saturation: 0.75 (higher saturation indicates rich image colors); eye openness: 0.85 (indicating that the eyes are fully open).
[0042] The predetermined emotional dimension refers to the emotional standard framework that has been set in emotional analysis, including emotional valence, arousal, and dominance. The predetermined emotional dimension is predefined and stored in the emotional analysis mechanism. When the emotional analysis of the image sequence corresponding to each shot in the source video is needed, the predefined emotional dimensions are first accessed. According to the feature parameter set corresponding to the image sequence, the coefficient of each emotional dimension is calculated to quantitatively describe the emotional state in the image. The emotional analysis mechanism is used to analyze and understand the emotional content in images, videos or other data. Emotional valence is an indicator in the emotional dimension, used to measure the polarity of emotion, i.e., the positive or negative degree of emotion. Generally, the emotional valence value is between -1 and +1, with negative values indicating negative emotions and positive values indicating positive emotions. Arousal measures the intensity of emotion, i.e., whether the emotion is calm or excited, usually between -1 and +1, with higher arousal values indicating excitement or happiness, and lower arousal values indicating calmness or tranquility. Dominance is a dimension that measures the control in an individual's emotional perception, usually measuring a person's sense of control or being controlled, usually between -1 and +1, with positive values indicating control or dominance of emotion, and negative values indicating passive or limited emotion.
[0043] Any one of the predetermined emotional dimensions is selected as the first emotional dimension, and the first feature parameter set is analyzed for emotion through the first emotional dimension, to obtain a dimension coefficient of the corresponding emotional dimension (such as valence, arousal, or dominance). The first dimension coefficient refers to a value related to the first emotional dimension calculated from the first feature parameter set in the process of emotional analysis. Through processing of multiple features extracted from the first feature parameter set (such as facial expression, sound frequency, color, etc.), a value related to the selected first emotional dimension (such as emotional valence) is calculated. For example, the emotional valence of the image or audio is inferred in combination with facial expression features (such as raised eyebrows, raised corners of the mouth), audio features (such as tone, speed), and the first dimension coefficient of emotional valence is 0.8, indicating that the emotional tendency is positive.
[0044] The first dimension coefficient is combined with the first emotional dimension and added to the first emotional information. For the first target image, other emotional dimensions in the predetermined emotional dimensions are also analyzed for emotion to obtain corresponding dimension coefficients, which together constitute the first emotional information. The first emotional information is the feature parameters extracted from the first target image, and the corresponding three dimension coefficients are calculated through the predetermined emotional dimensions, thereby constructing a complete three-dimensional emotional vector, for example, emotional valence is 0.8 (emotional positive), arousal is 0.61 (high excitement), and dominance is 0.54 (moderate dominance). Through extraction of multi-dimensional features and calculation of emotional dimension coefficients, the emotional features in the video are accurately reflected, especially in interactive scenes such as customer service, which can evaluate emotional changes in real time. By simultaneously calculating emotional valence, arousal, and dominance, the emotional state is fully captured, not only focusing on the positive and negative polarity of emotion, but also considering the intensity and control of emotion.
[0045] S200: retrieving an auditory modality analysis strategy to analyze the source audio extracted from the source video to obtain auditory analysis information.
[0046] Further, the S200 of the present application includes:
[0047] S210: obtaining a sound signal of the source audio, and analyzing the sound signal to obtain a sound intensity time sequence; S220: randomly dividing the sound intensity time sequence to obtain a division result; S230: collecting a first time-domain feature parameter of a first time sequence in the division result; S240: collecting a first frequency-domain feature parameter of a first frequency spectrum obtained by performing fast Fourier transform on the first time sequence; S250: activating a sound emotion prediction model in the hearing modality analysis strategy to perform prediction analysis on the first time-domain feature parameter and the first frequency-domain feature parameter, to obtain first predicted emotion information; and S260: matching a first audio segment corresponding to the first time sequence in the source audio, and establishing the hearing analysis information based on a second mapping relationship between the first audio segment and the first predicted emotion information.
[0048] Specifically, the source audio in the source video is obtained, i.e., the source audio containing voice, background noise and other sound information extracted from the source video. The sound signal is extracted from the source audio, which is performed by audio sampling. The audio signal is divided into multiple small time segments, and the amplitude and frequency information of each segment are extracted. The sound signal is usually sampled at a fixed sampling rate. The sampling rate is generally 16 kHz or 44.1 kHz, i.e., 16000 or 44100 sampling points per second. After sampling, the sound signal is represented as a set of digitized amplitude data, which reflects the amplitude of the sound waveform at each sampling time.
[0049] The sound intensity time sequence is obtained by calculating the amplitude (i.e., amplitude) of the sound signal. The greater the amplitude, the stronger the sound, and the smaller the amplitude, the weaker the sound. The amplitude of the sound signal is usually fluctuating in time, and the sound intensity time sequence is the change of these amplitude values over time. For example, the audio signal is divided into multiple time windows, and the commonly used window length is 20-50 milliseconds (ms). The frequency component and amplitude of each time window are calculated to obtain the short-time sound intensity, thereby constructing the sound intensity time sequence. The sound intensity time sequence describes the change process of the audio signal in time, usually representing the change of sound intensity or amplitude over time, reflecting the change trend of the sound, such as the increase and decrease of volume, the ups and downs of audio, etc. For example, the root mean square energy of the sampling point amplitude of the first 25 ms window of a certain customer audio is 0.2 (normalized amplitude, range 0~1); the root mean square energy of the second window is 0.18, and the root mean square energy of the third window is 0.25. After 400 window calculations, a sound intensity time sequence with a length of 400 is obtained, reflecting the dynamic change of the volume in 10 seconds.
[0050] The sound intensity time sequence is randomly divided into multiple time segments, and the sound intensity time sequence data is randomly divided into multiple time segments. The purpose is to analyze the sound signal locally and avoid deriving emotion prediction from the entire time sequence, thereby improving the generalization ability of the model. The first time sequence is extracted from the segmentation result, and the first time domain feature parameters are collected, including the average amplitude, maximum amplitude, variance, energy, zero-crossing rate, etc., which are used to describe the properties of the sound signal in the time domain. For example, assuming that the overall sound intensity time sequence length is 2800 points, the first time sequence length is 450 points after random division, which constitutes the first time sequence, and the corresponding first time domain feature parameters are collected to obtain the average amplitude 0.22, the variance 0.015, the maximum value 0.55, the zero-crossing rate 0.08, and the short-time energy 0.048.
[0051] The first time domain feature parameters are subjected to fast Fourier transform processing, the time window data (original sound waveform or amplitude sequence) of the sound signal corresponding to the first time sequence is taken, the fast Fourier transform algorithm is applied to the time sequence, the time domain signal is converted to the frequency domain, and the first frequency spectrum is obtained. The first frequency domain feature parameters of the first frequency spectrum, such as the dominant frequency, spectral centroid, bandwidth, frequency energy distribution, etc., are collected, which reflect the frequency characteristics of the signal. The frequency spectrum is an array of frequency and corresponding amplitude, which is expressed as the sound energy size of each frequency band. For example, the 1st element in the frequency spectrum corresponds to 0Hz (direct current component), the 10th element corresponds to about 355Hz (10x35.56Hz), and so on.
[0052] The sound emotion prediction model in the auditory modality analysis strategy is activated, and the first time domain feature parameters and the first frequency domain feature parameters are input into the trained sound emotion prediction model for prediction analysis. The first predicted emotion information is obtained, that is, for each input audio segment, the model gives the predicted values of the three emotion dimensions. The first predicted emotion information is the emotion prediction result generated by the sound emotion prediction model according to the input time domain features and frequency domain features, including the values of emotion valence, arousal and dominance.
[0053] In the source audio, according to the setting of the time window (such as every 2 seconds as a time sequence), the first audio segment corresponding to the first time sequence is matched, and the second mapping relationship between the first audio segment and the first predicted emotion information is established. Through the second mapping relationship, the predicted emotion information is associated with the actual audio segment. For example, the emotion valence of the audio segment from 10 seconds to 12 seconds is 0.7, the arousal is 0.5, and the dominance is 0.3. For other time sequences in the source audio segmentation result, repeat the above steps to obtain the predicted emotion information corresponding to each audio segment and determine the mapping relationship. The mapping relationship between all audio segments and their corresponding predicted emotion information is combined to form the auditory analysis information. The auditory analysis information is a set of information obtained by emotion prediction and analysis of the audio, including the emotion dimension values of each audio segment.
[0054] By randomly dividing the sound intensity time sequence, the original data is divided into multiple segments, which is helpful for extracting and analyzing the characteristics of the sound signal. By collecting time domain feature parameters and frequency domain feature parameters, the time domain and frequency domain properties of the sound signal are determined, and the sound emotion prediction model is used to predict the emotional state in the sound signal, which helps to analyze the audio content more deeply.
[0055] Further, the present application further comprises the following steps:
[0056] The sound emotion prediction model comprises three prediction channels of the predetermined emotional dimensions, and the predetermined emotional dimensions comprise emotional valence, arousal and dominance.
[0057] A sound emotion dictionary is constructed, and a first training data set of the first emotional dimension in the predetermined emotional dimensions is formed based on the sound emotion dictionary; supervised learning is performed on the first training data set to obtain a first prediction channel, and the sound emotion prediction model is formed; wherein the first training data set comprises first training time domain feature parameters, first training frequency domain feature parameters and first training emotional information.
[0058] Specifically, a large number of sound samples with emotional labels are collected, including customer service recordings, public emotional speech databases, each sound sample is manually or semi-automatically labeled to obtain its numerical labels in emotional valence, arousal and dominance (usually continuous values or discrete levels between 0 and 1). The sound emotion dictionary simultaneously saves the time domain and frequency domain feature parameters corresponding to the sample to form a complete training sample.
[0059] For the first emotional dimension in the predetermined emotional dimensions, the time domain features, frequency domain features and corresponding emotional valence related to the first emotional dimension are extracted from the sound emotion dictionary to obtain the first training time domain feature parameters, the first training frequency domain feature parameters and the first training emotional information, which constitute the first training data set. Data preprocessing is performed on the first training data set, including data cleaning (eliminating missing or abnormal samples) and normalization processing (standardizing the time domain and frequency domain features to the range of 0~1 to reduce the influence of different dimensions). Select a suitable machine learning algorithm, such as a double-layer LSTM model. The structure of the double-layer LSTM model is constructed, the input layer is a 50-dimensional feature vector, the first LSTM layer is 128 units with Dropout (0.2), the second LSTM layer is 64 units, and the full connection layer is one neuron to output the emotional valence prediction.
[0060] The model training is performed using the first training data set, and a loss function such as mean square error (MSE) is used to measure the prediction error. The model parameters are adjusted through the back propagation algorithm to gradually reduce the loss. During the training, all samples of the first training data set are iterated in each training period. The model performance is evaluated on the validation set after each training period, and the hyperparameters are adjusted. If the validation set loss does not improve for 5 consecutive training periods, the training is stopped.
[0061] For the other two prediction channels in the sound emotion prediction model, the above steps are repeated to obtain the second and third prediction channels, which form the sound emotion prediction model. The sound emotion prediction model can predict emotion dimensions (such as valence, arousal, and dominance) based on the features of the audio signal. By constructing a sound emotion dictionary, a training data set related to emotion is established, which helps to train the sound emotion prediction model. Using a supervised learning method, the model is trained based on known input and output data, so that it can predict unknown data. The sound emotion prediction model can provide an understanding of the emotional level of audio content, which helps to improve the accuracy and depth of multi-modal semantic analysis.
[0062] S300: Perform stage-by-stage comparative analysis on the visual analysis information and the auditory analysis information to obtain an emotional contrast result.
[0063] Specifically, the visual analysis information and the auditory analysis information are subjected to stage-by-stage comparative analysis, i.e., the visual analysis information and the auditory analysis information are compared and analyzed in different time periods to evaluate their consistency or difference in emotion. The purpose of comparative analysis is to detect whether the visual and auditory information is consistent, thereby enhancing the accuracy of emotion recognition and avoiding the bias of a single modality.
[0064] The visual analysis information and the auditory analysis information are aligned in chronological order, and the valence, arousal, and dominance of different time periods are compared. For each time period, the Pearson correlation coefficient is used to calculate the similarity or difference of the valence, arousal, and dominance of the visual and auditory information. If the Pearson correlation coefficient is high (close to 1), it means that the visual and auditory information is consistent in that time period; if the Pearson correlation coefficient is low (close to 0 or negative), it means that the visual and auditory information may conflict, and further adjustment or weighting processing is needed.
[0065] After stage-by-stage comparative analysis, the emotional contrast results of each time period are output. For example, in the time period 10s-12s, the visual valence = 0.7, the auditory valence = 0.68, which is high consistency; the visual arousal = 0.8, the auditory arousal = 0.75, which is high consistency; the visual dominance = 0.6, the auditory dominance = 0.65, which is high consistency. In the time period 20s-25s, the visual valence = -0.2, the auditory valence = -0.1, which is lower consistency; the visual arousal = 0.6, the auditory arousal = 0.3, which is medium consistency; the visual dominance = 0.5, the auditory dominance = 0.3, which is lower consistency.
[0066] Through stage-by-stage comparative analysis, the overall emotional contrast results are obtained. The emotional contrast results reveal the differences or consistency of the video and the audio in emotional expression. For example, if the emotions of the video and the audio are inconsistent in some time periods, it may indicate that the customer expresses different emotions in the video than in the voice (such as hiding true emotions through facial expressions, or inconsistent tone and facial expressions). According to the contrast results, corresponding adjustments are made, such as strengthening the processing, weighting influence, etc. of the audio or video signal. By comparing and analyzing the visual and auditory analysis information, the consistency of different modalities in the emotional dimension is determined, and the emotional state of the user is comprehensively understood.
[0067] S400: If the emotional contrast result reaches a predetermined contrast constraint, analyze the modal analysis database to obtain real-time modal weight distribution.
[0068] Further, the S400 of the present application comprises:
[0069] S410: sampling in the modal analysis database to obtain modal analysis samples; S420: manually reviewing the modal analysis samples to obtain review results; S430: analyzing the review results to obtain modal analysis accuracy, wherein the modal analysis accuracy includes visual analysis accuracy and auditory analysis accuracy; S440: performing variation weighting analysis on the visual analysis accuracy and the auditory analysis accuracy to obtain the real-time modal weight distribution.
[0070] Specifically, the predetermined contrast constraint is a set emotional consistency threshold or rule. If the predetermined contrast constraint is reached, it means that the emotional information of the two modalities is close enough, and the analysis can continue. The modal analysis database is a stored multi-modal emotional analysis result library, which contains a large number of historical sample data with visual analysis information, auditory analysis information, original segments, time labels, etc.
[0071] When the emotional contrast results of the visual and auditory modalities reach the predetermined contrast constraint, it indicates that the consistency of the multi-modal emotional information has reached the trust standard. The modal analysis samples are obtained by sampling from the modal analysis database. The modal analysis samples include visual analysis samples and auditory analysis samples, which are used for artificial emotional review. The artificial emotional review is performed on the modal analysis samples, and the subjective evaluation score of the sampled emotional analysis samples is obtained by experts as a true emotional label reference. The modal analysis accuracy is obtained by analyzing the review results using mean square error (MSE). The consistency between the results of the visual and auditory modalities in the emotional analysis and the artificial review results is analyzed to obtain the modal analysis accuracy. The higher the accuracy, the closer the emotional prediction of the modal analysis strategy under the modality to the true emotion annotated by artificial labeling. The visual analysis accuracy is the accuracy of the visual modality (such as facial expression, image analysis) emotional analysis result. The auditory analysis accuracy is the accuracy of the auditory modality (such as speech analysis, audio feature extraction) emotional analysis result.
[0072] The weight is calculated according to the fluctuation of the visual analysis accuracy and the auditory analysis accuracy, that is, the standard deviation of the accuracy is used as the weight adjustment basis. The modality with smaller standard deviation (i.e., the modality with more stable accuracy) is given a higher weight; the modality with larger standard deviation (i.e., the modality with greater fluctuation of accuracy) is given a lower weight. For example, the average error of the visual modality prediction and the artificial is: valence MAE=0.07, arousal MAE=0.10, dominance MAE=0.06, the visual MAE is calculated as 0.0767, and the visual analysis accuracy is 1-0.0767=92.33%; the average error of the visual modality prediction and the artificial is: valence MAE=0.08, arousal MAE=0.11, dominance MAE=0.09, the auditory MAE is calculated as 0.0933, and the auditory analysis accuracy is 90.67%. The visual standard deviation is 0.015, the auditory standard deviation is 0.03, the stability index is calculated as: visual stability index 61.55, auditory stability index 30.22, and thus the real-time modality weight distribution obtained is: visual modality weight 67%; auditory modality weight 33%. According to the current model accuracy and its stability, the modality weight is adaptively distributed to avoid large errors of a certain modality leading to overall misjudgment. The weight adaptive distribution mechanism avoids the fuzzy judgment caused by the average processing mode, and has more judgment power in multi-source information fusion.
[0073] S500: Obtain text analysis information of the source text, and combine the real-time modality weight distribution to fuse the text analysis information, the visual analysis information, and the auditory analysis information to obtain a semantic analysis result.
[0074] Further, the S500 of the present application comprises:
[0075] S510: performing cleaning and word segmentation processing on the source text to obtain a processing result; S520: performing feature analysis on the processing result based on the bag-of-words model principle to obtain a feature vector; S530: performing sentiment analysis on the feature vector in combination with a vocabulary sentiment dictionary to obtain the text analysis information.
[0076] Further, the present application further comprises the following steps:
[0077] S531: extracting any vocabulary in the feature vector; S532: matching any score of the any vocabulary in the vocabulary sentiment dictionary, wherein the any score comprises three scores of the predetermined sentiment dimensions; S533: obtaining any frequency of the any vocabulary according to the feature vector, and obtaining any sentiment score of the any vocabulary in combination with the three scores; S534: taking the mean of the any sentiment score to form the text analysis information.
[0078] Specifically, the source text is subjected to cleaning and word segmentation processing to remove irrelevant characters, punctuation marks and stop words, and the text is segmented into words or vocabulary units to obtain a processing result. The text is preprocessed to remove noise (such as punctuation marks, stop words, etc.), and word segmentation is performed to divide the text into individual words. The source text is the original input text data, such as the conversation content between the customer and the customer service obtained from the customer service conversation, or the text extracted from the video subtitles. Cleaning includes removing all unnecessary symbols, punctuation marks, stop words (such as de, le, etc.). The text is divided into independent words using a word segmentation algorithm, including rule-based word segmentation (such as regular expressions) or statistical-based word segmentation (such as jieba segmentation, NLP toolkit, etc.).
[0079] The bag-of-words model is a text representation method that ignores the order of words and only considers the frequency of each word in the text. Feature analysis is to convert the processed text into a high-dimensional vector, where each dimension corresponds to a word, and the value is the frequency of the word in the text. Based on the bag-of-words model principle, the processing result is subjected to feature analysis to obtain a feature vector. In this process, each word is a feature, and each dimension of the feature vector corresponds to a word, and its value represents the frequency of the word in the text. Specifically, all the words of the text are collected to construct a vocabulary table (dictionary) containing all the independent words. For each text, the number of occurrences of each word in the vocabulary table is counted. The word frequency is converted into a vector, and the vector length is equal to the size of the vocabulary table. According to the word frequency statistics, the text is converted into a feature vector using the bag-of-words model. The length of the feature vector is equal to the size of the vocabulary table. The value of each position corresponds to the number of occurrences of the corresponding word in the text.
[0080] Extract any vocabulary in the feature vector, match any score of any vocabulary in the vocabulary light point dictionary, including three scores of the predetermined sentiment dimension, namely the sentiment valence score, the arousal score and the dominance score. The sentiment score of each vocabulary in the sentiment dictionary is usually composed of three sentiment dimensions: valence, arousal and dominance. The vocabulary sentiment dictionary is a dictionary containing vocabulary and its corresponding sentiment score. Each vocabulary is associated with multiple sentiment dimensions, such as valence, arousal and dominance.
[0081] According to the frequency of any vocabulary extracted from the feature vector, the sentiment score of any vocabulary is calculated in combination with the three scores. The frequency of any vocabulary is the number of times it appears. The sentiment score of each vocabulary is calculated based on the frequency of the vocabulary in the text and its score in the sentiment dictionary. For example, the sentiment score of the customer service vocabulary is (valence 0.7, arousal 0.4, dominance 0.5), and the frequency is 1. According to the frequency and sentiment score of the customer service, the sentiment score of the vocabulary is calculated as (0.7, 0.4, 0.5).
[0082] The mean value of the arbitrary sentiment score is calculated to form the text analysis information. For each vocabulary in the feature vector, repeat the above steps to form the text analysis information. According to the real-time modal weight distribution, the text analysis information, visual analysis information and auditory analysis information are fused, that is, according to the analysis accuracy of different modalities (text, vision, hearing) at a certain moment, the weight of each modality is distributed by weighting, and the weighted average of the text analysis information, visual analysis information and auditory analysis information after adjusting the weight is used to obtain the semantic analysis result. Through real-time modal weight distribution, the contribution of each modality in semantic analysis can be dynamically adjusted according to the analysis accuracy of each modality, ensuring the reliability of the final result. The fusion of sentiment information of text, visual and auditory modalities can comprehensively reflect the sentiment state of the input data, making the semantic analysis more in-depth and diverse.
[0083] In summary, the cross-modal semantic analysis method provided by the present application has the following beneficial effects:
[0084] The source video is parsed by calling a visual modal analysis strategy to obtain visual analysis information; the source audio extracted from the source video is parsed by calling an auditory modal analysis strategy to obtain auditory analysis information; the visual analysis information and the auditory analysis information are subjected to stage-by-stage comparative analysis to obtain a sentiment comparison result; if the sentiment comparison result reaches a predetermined comparison constraint, real-time modal weight distribution is obtained by analyzing a modal analysis database; text analysis information of source text is obtained, and the text analysis information, the visual analysis information and the auditory analysis information are fused in combination with the real-time modal weight distribution to obtain a semantic analysis result. That is, visual analysis information is extracted by parsing visual content in a source video; auditory analysis information is obtained by parsing source audio in the source video; a sentiment comparison result is obtained by stage-by-stage comparative analysis of the visual analysis information and the auditory analysis information; the weight of each modal is dynamically adjusted by analyzing the sentiment comparison result; the analysis information of text, vision and hearing is fused according to real-time modal weight distribution to obtain a semantic analysis result, thereby improving the accuracy and adaptability of cross-modal semantic analysis.
[0085] The above description of disclosed embodiments enables one skilled in the art to make or use the application. Numerous modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0086] Obviously, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A cross-modal semantic parsing method, characterized in that, The method comprises the following steps: a visual modality analysis strategy is invoked to analyze the source video to obtain visual analysis information; an auditory modality analysis strategy is invoked to analyze the source audio extracted from the source video to obtain auditory analysis information; the visual analysis information and the auditory analysis information are subjected to stage-by-stage comparative analysis to obtain a sentiment comparison result; if the sentiment comparison result reaches a predetermined comparison constraint, real-time modality weight distribution is obtained by analyzing a modality analysis database; text analysis information of source text is obtained, and the text analysis information, the visual analysis information and the auditory analysis information are fused in combination with the real-time modality weight distribution to obtain semantic analysis results; the method comprises the following steps: the source video is subjected to dynamic image compression processing to obtain a compressed video; the compressed video is subjected to detection analysis according to a shot detection mechanism in the visual modality analysis strategy to obtain a detection result; a first image sequence corresponding to a first shot in the detection result is extracted, and the first image sequence is subjected to enhancement fusion processing to obtain a first target image; the first target image is subjected to sentiment analysis according to a sentiment analysis mechanism in the visual modality analysis strategy to obtain first sentiment information; a first video segment corresponding to the first shot in the source video is matched, and the visual analysis information is established based on a first mapping relationship between the first video segment and the first sentiment information; the method comprises the following steps: a sound signal of the source audio is obtained, and the sound signal is analyzed to obtain a sound intensity time sequence; the sound intensity time sequence is subjected to random segmentation to obtain a segmentation result; first time-domain feature parameters of a first time sequence in the segmentation result are collected; first frequency-domain feature parameters of a first frequency spectrum obtained by fast Fourier transform processing of the first time sequence are collected; a sound sentiment prediction model in the auditory modality analysis strategy is activated to perform prediction analysis on the first time-domain feature parameters and the first frequency-domain feature parameters to obtain first predicted sentiment information; a first audio segment corresponding to the first time sequence in the source audio is matched, and the auditory analysis information is established based on a second mapping relationship between the first audio segment and the first predicted sentiment information.
2. The cross-modal semantic parsing method of claim 1, wherein, the method comprises the following steps: adjacent image groups of the compressed video are obtained; adjacent displacement values of a first image and a second image in the adjacent image groups are calculated according to the shot detection mechanism; if the adjacent displacement values do not conform to a predetermined displacement limit value, the first image and the second image are different shots, and the detection result is formed.
3. The cross-modal semantic parsing method of claim 1, wherein, the method comprises the following steps: a first feature parameter set of the first target image is collected in multiple dimensions; extracting a first emotional dimension in a predetermined emotional dimension, wherein the predetermined emotional dimension is stored in the emotional analysis mechanism; obtaining a first dimension coefficient of the first emotional dimension in combination with the first feature parameter set; constructing the first emotional information based on the first dimension coefficient.
4. The cross-modal semantic parsing method of claim 3, wherein, The sound emotion prediction model includes three prediction channels of the predetermined emotional dimension, and the predetermined emotional dimension includes emotional valence, arousal and dominance.
5. The cross-modal semantic parsing method of claim 4, wherein, The construction of the sound emotion prediction model includes: constructing a sound emotion dictionary, and constructing a first training data set of the first emotional dimension in the predetermined emotional dimension based on the sound emotion dictionary; performing supervised learning on the first training data set to obtain a first prediction channel, and composing the sound emotion prediction model; wherein the first training data set includes first training time domain feature parameters, first training frequency domain feature parameters and first training emotional information.
6. The cross-modal semantic parsing method of claim 1, wherein, If the emotional comparison result reaches a predetermined comparison constraint, analyze the modal analysis database to obtain a real-time modal weight distribution, including: sampling in the modal analysis database to obtain a modal analysis sample; performing artificial emotional review on the modal analysis sample to obtain a review result; analyzing the review result to obtain a modal analysis accuracy, wherein the modal analysis accuracy includes a visual analysis accuracy and an auditory analysis accuracy; performing variation weighting analysis on the visual analysis accuracy and the auditory analysis accuracy to obtain the real-time modal weight distribution.
7. The cross-modal semantic parsing method of claim 3, wherein, obtaining the text analysis information of the source text, including: performing cleaning and word segmentation processing on the source text to obtain a processing result; performing feature analysis on the processing result based on the bag-of-words model principle to obtain a feature vector; performing emotional analysis on the feature vector in combination with the lexical emotion dictionary to obtain the text analysis information.
8. The cross-modal semantic parsing method of claim 7, wherein, Performing emotional analysis on the feature vector in combination with the lexical emotion dictionary to obtain the text analysis information, including: extracting any lexical in the feature vector; matching to obtain any score of the any lexical in the lexical emotion dictionary, wherein the any score includes three scores of the predetermined emotional dimension; obtaining any frequency of the any lexical according to the feature vector, and obtaining any emotional score of the any lexical in combination with the three scores; taking the mean of the any emotional score to compose the text analysis information.
Citation Information
Patent Citations
Audio-visual video analysis device and method based on multi-scale semantic network
CN114519809A
Implicit sentiment analysis method based on vision and speech enhancement
CN118395233A