Data processing method and device, electronic equipment and computer readable storage medium

By combining voice and facial images for fine-grained emotion recognition and adaptive subtitle position adjustment, the problem of lack of fine-grained emotion and position adaptation in smart glasses subtitles has been solved, improving the reading experience and data compliance.

CN121262437APending Publication Date: 2026-01-02WUHAN BOBODONG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511276403.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing smart glasses generate subtitles that lack fine-grained emotion recognition and cannot adaptively adjust the subtitle position, resulting in a poor reading experience in noisy environments and multi-person conversation scenarios.

Method used

By combining speech data and facial images to predict emotion categories, fine-grained emotion captions are generated. Eye-tracking data and location information are used to adaptively adjust the caption position, and differential privacy federated learning is used to optimize the emotion recognition model.

Benefits of technology

It improved the accuracy of sentiment prediction results, reduced eye movement load, enhanced the subtitle reading experience in multi-person dialogue scenarios, and met data compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121262437A_ABST
    Figure CN121262437A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a computer readable storage medium. The method comprises the steps of obtaining voice data and a face image corresponding to a first object; when determining that an emotion prediction condition is satisfied based on the voice data and the face image, performing emotion category prediction based on at least one of the voice data and the face image to obtain an emotion prediction result; generating text data based on the voice data, and generating subtitle data based on the emotion prediction result and the text data; acquiring eye movement data corresponding to the second object and position information of the first object relative to the second object; and determining a target position of the subtitle data based on the position information and the eye movement data, and displaying the subtitle data at the target position. According to the invention, the fine-grained emotion corresponding to the subtitles can be presented, and the positions of the subtitles can be adaptively adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a data processing method and device, electronic equipment and computer readable storage medium. BACKGROUND

[0002] In daily communication, business meetings, education training and other scenarios, a user can need to obtain text information of dialogue content in real time. Especially in a noisy environment, or when the user wants to avoid disturbing others, smart glasses generating subtitles can provide a silent way of information acquisition. For users communicating across countries or learning foreign languages, smart glasses can translate voice content into subtitles in a target language in real time, helping users overcome language barriers and improve communication efficiency. Hearing-impaired people often face obstacles in daily communication, and smart glasses generating subtitles can provide real-time text information to enhance communication ability and participation.

[0003] In related technologies, the subtitles generated by smart glasses lack fine-grained emotions, and the subtitles are displayed in a fixed position on the screen, which cannot realize adaptive adjustment of the subtitle position. SUMMARY

[0004] The embodiments of the present application provide a data processing method and device, electronic equipment and computer readable storage medium, which can present fine-grained emotions corresponding to subtitles and adaptively adjust the position of the subtitles.

[0005] The technical solutions of the embodiments of the present application are implemented as follows:

[0006] The embodiments of the present application provide a data processing method, which comprises the following steps:

[0007] Obtain voice data and facial images corresponding to a first object;

[0008] When it is determined that an emotion prediction condition is met based on the voice data and the facial images, perform emotion category prediction based on at least one of the voice data and the facial images to obtain an emotion prediction result;

[0009] Generate text data based on the voice data, and generate subtitle data based on the emotion prediction result and the text data;

[0010] Obtain eye movement data corresponding to a second object, and position information of the first object relative to the second object;

[0011] Determine a target position of the subtitle data based on the position information and the eye movement data, and display the subtitle data at the target position.

[0012] The embodiments of the present application provide a data processing device, which comprises:

[0013] The first obtaining module is configured to obtain voice data and a facial image corresponding to a first object.

[0014] The emotion prediction module is configured to, when it is determined that an emotion prediction condition is met based on the voice data and the facial image, perform emotion category prediction based on at least one of the voice data and the facial image, to obtain an emotion prediction result.

[0015] The data generation module is configured to generate text data based on the voice data, and generate subtitle data based on the emotion prediction result and the text data.

[0016] The second obtaining module is configured to obtain eye movement data corresponding to a second object, and position information of the first object relative to the second object.

[0017] The position determination module is configured to determine a target position of the subtitle data based on the position information and the eye movement data, and display the subtitle data at the target position.

[0018] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:

[0019] The memory is configured to store computer executable instructions or computer programs.

[0020] The processor is configured to execute the computer executable instructions or computer programs stored in the memory, to implement the data processing method provided in the embodiments of the present application.

[0021] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores computer executable instructions or computer programs, and is configured to be executed by a processor to implement the data processing method provided in the embodiments of the present application.

[0022] A computer program product is provided in an embodiment of the present application, and the computer program product comprises computer executable instructions or computer programs, and the computer executable instructions or computer programs are executed by a processor to implement the data processing method provided in the embodiments of the present application.

[0023] The embodiments of the present application have the following beneficial effects:

[0024] By applying the embodiment of the application, the speech data and the face image corresponding to the first object are acquired, when it is determined that the emotion prediction condition is met based on the speech data and the face image, the emotion category prediction is performed based on at least one of the speech data and the face image, and the emotion prediction result is obtained, the emotion category of the subtitle can be predicted in combination with the speech data and the face image, the fine-grained emotion of the subtitle is realized, and the accuracy of the emotion prediction result is improved, then the text data is generated based on the speech data, and the subtitle data is generated based on the emotion prediction result and the text data, the fine-grained emotion of the subtitle can be presented, the eye movement data corresponding to the second object and the position information of the first object relative to the second object are further acquired, and then the target position of the subtitle data is determined based on the position information and the eye movement data, and the subtitle data is displayed at the target position. In this way, the eye movement tracking information and the speaker object positioning information are combined to realize the adaptive adjustment of the subtitle position, the visual migration load is reduced, and thus the subtitle reading experience in the multi-person dialogue scene is improved. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 is an application mode schematic diagram of the data processing method provided by the embodiment of the application;

[0026] Figure 2 is a structural schematic diagram of an electronic device provided by the embodiment of the application;

[0027] Figure 3A is a first flow schematic diagram of the data processing method provided by the embodiment of the application;

[0028] Figure 3B is a second flow schematic diagram of the data processing method provided by the embodiment of the application;

[0029] Figure 3C is a third flow schematic diagram of the data processing method provided by the embodiment of the application.

[0030] It should be noted that the "first" and "second" above are only used to distinguish different schemes, and do not represent the degree of superiority or priority in the implementation process. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical scheme and advantages of the application more clear, the application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the application, and all other embodiments obtained by a person skilled in the art without creative labor are within the scope of protection of the application.

[0032] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" can be the same subset or different subsets as each other and can be combined with each other as long as there is no conflict.

[0033] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects. It is understood that "first\second\third" can be interchanged in a specific order or sequence as long as it is allowed, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein.

[0034] The related data collection processing in the embodiments of the application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.

[0035] In the embodiments of the application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.

[0036] Unless otherwise defined, all technical and scientific terms used in the embodiments of the application have the same meanings as those commonly understood by one of ordinary skill in the art. The terms used in the embodiments of the application are only for the purpose of describing the embodiments of the application and are not intended to limit the application.

[0037] Before the embodiments of the application are further described in detail, the terms and phrases involved in the embodiments of the application are explained, and the terms and phrases involved in the embodiments of the application are applicable to the following explanations.

[0038] 1) Graph Attention Network (GAT): It is a graph neural network model based on attention mechanism, which efficiently processes graph structured data by dynamically learning the importance weight between nodes.

[0039] 2) Differential privacy federated learning: It is a technology that combines differential privacy and federated learning, aiming to protect data privacy while achieving efficient distributed model training.

[0040] The embodiment of the present application provides a data processing method and device, electronic equipment and computer readable storage medium, which can present fine-grained emotions corresponding to subtitles and adaptively adjust the positions of the subtitles.

[0041] The following describes an exemplary application of the electronic equipment provided by the embodiment of the present application. The electronic equipment provided by the embodiment of the present application can be implemented as various types of terminals such as smart wearable devices, notebook computers, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart televisions, vehicle-mounted terminals, augmented reality (AR) glasses, smart glasses, and the like. In the following, an exemplary application when the electronic equipment is implemented as smart glasses will be described.

[0042] Referring to Figure 1 , Figure 1 is an application mode schematic diagram of the data processing method provided by the embodiment of the present application, and Figure 1 involve a first object 200, a second object 300, and a terminal 400, and Figure 1 in the embodiment, the terminal 400 is exemplarily shown as smart glasses, and the smart glasses are worn on the second object.

[0043] In the process of generating subtitles, the terminal 400 acquires speech data and facial images corresponding to the first object 200. When it is determined that an emotion prediction condition is met based on the speech data and the facial images, an emotion category prediction is performed based on at least one of the speech data and the facial images to obtain an emotion prediction result. Text data is generated based on the speech data, and subtitle data is generated based on the emotion prediction result and the text data. The terminal 400 acquires eye movement data corresponding to the second object 300 and position information of the first object 200 relative to the second object 300. Based on the position information and the eye movement data, a target position of the subtitle data is determined, and the subtitle data is displayed at the target position. The terminal 400 can be a smart phone, augmented reality (AR) glasses, and the like used in online meeting, game live broadcast, online education, court record, real-time conversation, barrier-free cinema, and the like. Exemplarily, in an online meeting scenario, the terminal 400 moves the determined subtitle data to the corresponding target position of the video window of each participant based on the determined subtitle data and the target position of the subtitle data. Or in a barrier-free cinema scenario, the terminal 400 moves the determined subtitle data to the corresponding target position in the movie character area based on the determined subtitle data and the target position of the subtitle data.

[0044] Referring to Figure 2 , Figure 2 is a structural schematic diagram of the electronic equipment provided by the embodiment of the present application, and the electronic equipment can be a terminal or a server, Figure 2The illustrated electronic device includes at least one processor 410, memory 450, at least one network interface 420. The various components of the electronic device are coupled together by a bus system 440, which can include a power bus, a control bus and a status signal bus. For the sake of clarity, the various buses are illustrated in FIG. 4 as the bus system 440. The use of the term "bus" in this description should be taken in an abstract, generic sense, and not in a highly-specific, electrical-engineering sense. Figure 2

[0045] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general purpose processor, a Digital Signal Processor (DSP), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or the like. The general purpose processor can be a microprocessor, or any conventional processor, etc.

[0046] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 optionally includes one or more storage devices remotely located from the processor 410.

[0047] The memory 450 includes volatile memory or non-volatile memory, or both. Non-volatile memory can be read only memory (ROM), volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0048] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are described below.

[0049] The operating system 451 includes a system program for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks.

[0050] The network communication module 452 is used to communicate with other electronic devices via one or more (wired or wireless) network interfaces 420, examples of which include Bluetooth, Wireless Fidelity (WiFi), Universal Serial Bus (USB), etc.

[0051] ​In some embodiments, the device provided by the embodiments of the present application can be implemented in software, Figure 2 The data processing device 455 stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: a first acquisition module 4551, an emotion prediction module 4552, a data generation module 4553, a second acquisition module 4554, and a position determination module 4555. These modules are logical, and thus can be combined or further split according to the implemented functions. The functions of the individual modules will be described below.

[0052] The data processing method provided by the embodiments of the present application will be described in conjunction with exemplary applications and implementations of the terminal provided by the embodiments of the present application.

[0053] Next, the data processing method provided by the embodiments of the present application will be described. To facilitate understanding of the data processing method provided by the embodiments of the present application, the application in the emotion self-adaptive real-time subtitle generation and layout scenario is taken as an example for description.

[0054] As described above, the electronic device implementing the data processing method of the embodiments of the present application can be a terminal. Referring to Figure 3A Figure 3A is a first flowchart of the data processing method provided by the embodiments of the present application, which will be described in conjunction with Figure 3A the steps shown.

[0055] In step 301, voice data and facial images corresponding to a first object are acquired.

[0056] Here, the first object is an object initiating a real-time conversation. The terminal can acquire the voice data corresponding to the first object through a built-in microphone array at a frequency of 16 kHz and a frame rate of 20 milliseconds, and can acquire the facial images corresponding to the first object through a front camera at a rate of 30 frames per second.

[0057] In step 302, when it is determined that the emotion prediction condition is met based on the voice data and the facial images, an emotion category prediction is performed based on at least one of the voice data and the facial images to obtain an emotion prediction result.

[0058] Here, the signal-to-noise ratio of the voice data and the occlusion rate of the facial images are used to determine whether the emotion prediction condition is met. When it is determined that the emotion prediction condition is met based on the signal-to-noise ratio of the voice data and the occlusion rate of the facial images, an emotion category prediction is performed based on at least one of the voice data and the facial images to obtain an emotion prediction result. The emotion category includes 12 categories such as joy, calmness, surprise, sadness, anger, disgust, fear, doubt, irony, embarrassment, fatigue, and excitement.

[0059] ​In some embodiments, the determination of the satisfaction of the emotion prediction condition can be achieved by determining a signal-to-noise ratio of the voice data and an occlusion rate of the facial image; and determining that the emotion prediction condition is satisfied when the signal-to-noise ratio of the voice data is greater than or equal to a first preset threshold or the occlusion rate of the facial image is less than or equal to a second preset threshold.

[0060] Here, the signal-to-noise ratio (SNR) refers to the ratio of the effective voice signal to the noise signal in the voice data, reflecting the intelligibility and intelligibility of the voice data. Voice data with high signal-to-noise ratio is easier to recognize and understand, which helps to improve the accuracy of voice recognition. Voice data with low signal-to-noise ratio may result in inaccurate voice recognition results, affecting the reliability of subsequent voice data processing. The voice activity detection technology is used to extract the effective voice signal from the voice data. The frequency domain analysis method is used to separate the noise signal from the voice data. The effective voice signal power and the noise signal power are calculated, and the power is usually represented as the square of the signal amplitude. The effective voice signal power and the noise signal power are subjected to inverse logarithmic operation to obtain the signal-to-noise ratio of the voice data.

[0061] The occlusion rate refers to the proportion of the face region in the facial image that is occluded or invisible by the occlusion object, which may include hair, hands, glasses, hats and other objects. High occlusion rate will reduce the detectability of the facial image and affect the accuracy of expression recognition. Low occlusion rate of the facial image helps to improve the accuracy of feature extraction and improve the processing effect of the facial image. The face detection algorithm (such as a deep learning model such as a multi-task convolutional neural network) is used to determine the face region in the facial image. The target detection algorithm is used to detect the occlusion object (such as hands, hair, etc.) in the image, to determine the total area of the face region and the occlusion area of the occlusion object in the face region, and to determine the ratio of the occlusion area to the total area of the face region as the occlusion rate of the facial image.

[0062] The first preset threshold is a preset signal-to-noise ratio threshold, and the second preset threshold is a preset occlusion rate threshold. When the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold or the occlusion rate of the facial image is less than or equal to the second preset threshold, it is determined that the emotion prediction condition is satisfied. For example, the first preset threshold is 10 decibels, and the second preset threshold is 40%. When the signal-to-noise ratio of the voice data is greater than or equal to 10 decibels or the occlusion rate of the facial image is less than or equal to 40%, it is determined that the emotion prediction condition is satisfied.

[0063] In some embodiments, referring to Figure 3B , Figure 3B is a second flowchart of the data processing method provided by the embodiments of the present application, Figure 3A The step 302 shown can be implemented by the steps 3021 to 3023 of the step 302, Figure 3B The step 302 shown can be implemented by the steps 3021 to 3023 of the step 302,

[0064] In step 3021, when the signal-to-noise ratio of the voice data is greater than or equal to a first preset threshold, and the occlusion rate of the face image is less than or equal to a second preset threshold, an emotion category prediction is performed based on the voice data and the face image to obtain an emotion prediction result.

[0065] Here, when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the face image is less than or equal to the second preset threshold, the voice data and the face image both satisfy the emotion prediction condition, and the emotion category prediction can be performed based on the voice data and the face image at the same time to obtain the emotion prediction result.

[0066] In some embodiments, the emotion category prediction based on the voice data and the face image to obtain the emotion prediction result can be achieved by the following steps: feature extraction is performed on the voice data and the face image respectively to obtain audio features and face features; the audio features and the face features are fused to obtain fused features; a graph attention network is used to perform emotion category prediction on the fused features to obtain an emotion category probability distribution; and emotion intensity mapping is performed based on the emotion category probability distribution to obtain the emotion prediction result.

[0067] Here, the voice data is feature-extracted by a convolutional neural network to obtain audio features, and the audio features include fundamental frequency, energy, zero-crossing rate, speech rate, spectral features, and 512-dimensional acoustic embedding, etc. rhythm features. The face image is feature-extracted by a lightweight face recognition model MobileFace2 to obtain face features, and the face features include 17 Action Unit (AU) intensity features.

[0068] The timestamps of the voice data and the face image are aligned by interpolation or synchronization technology, and a time window of 200 milliseconds is used. In a plurality of continuous time windows, the audio features are determined as rhythm nodes, and the face features are determined as face nodes. In the same time window, the mutual information between the rhythm nodes and the face nodes, i.e., the correlation between the rhythm nodes and the face nodes, is determined. The rhythm nodes and the face nodes with mutual information higher than a threshold are connected to realize the fusion of the audio features and the face features, and the fused features are obtained.

[0069] The fusion feature is input to a full connection layer of a graph attention network (GAT) to perform emotion category prediction on the fusion feature, to obtain emotion category probability distributions of 12 emotion categories, and based on the emotion category probability distributions, an emotion category with a maximum probability is determined as a target emotion category. Each emotion category corresponds to a preset emotion intensity, and the emotion intensity includes five levels of extremely negative, negative, neutral, positive, and extremely positive. Based on the preset emotion intensity corresponding to the target emotion category, emotion intensity mapping is performed on the target emotion category to obtain an emotion prediction result.

[0070] In the embodiments of the present application, emotion category prediction is performed based on voice data and facial images to obtain an emotion prediction result. The voice data can reflect voice information of a speaking object, and the facial images can provide visual emotional information. Multi-modal emotion category prediction is performed by combining the voice data and the facial images, thereby improving the accuracy of the emotion prediction result.

[0071] With reference to Figure 3B In step 3022, when the signal-to-noise ratio of the voice data is less than a first preset threshold, and the occlusion rate of the facial image is less than or equal to a second preset threshold, emotion category prediction is performed based on the facial image to obtain an emotion prediction result.

[0072] Here, when the signal-to-noise ratio of the voice data is less than the first preset threshold, and the occlusion rate of the facial image is less than or equal to the second preset threshold, that is, the facial image satisfies the emotion prediction condition, emotion category prediction can be performed based on the facial image to obtain an emotion prediction result.

[0073] In some embodiments, a facial key point detection algorithm is used to locate a facial region in the facial image. A feature extraction layer in a lightweight face recognition model MobileFace2 is used to perform feature extraction on the facial region to obtain facial features, which include 17 Action Unit (AU) intensity features. A full connection layer in the lightweight face recognition model is used to perform emotion category mapping on the facial features to obtain an emotion prediction result.

[0074] With reference to Figure 3B In step 3023, when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the facial image is greater than the second preset threshold, emotion category prediction is performed based on the voice data to obtain an emotion prediction result.

[0075] Here, when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the facial image is greater than the second preset threshold, that is, the voice data satisfies the emotion prediction condition, emotion category prediction can be performed based on the voice data to obtain an emotion prediction result.

[0076] In some embodiments, the voice data is divided into a plurality of time windows, and in each time window, a convolutional neural network in the emotion recognition model is used for feature extraction of the voice data to obtain audio features, including fundamental frequency, energy, zero-crossing rate, speech rate, spectral features, and 512-dimensional acoustic embedding, etc. rhythm features. Through the full connection layer in the emotion recognition model, the audio features are mapped to the emotion categories to obtain the emotion prediction result.

[0077] In the embodiments of the present application, the emotion prediction condition is determined based on the signal-to-noise ratio of the voice data and the occlusion rate of the face image, which can avoid the influence of low-quality voice data and face image on the accuracy of the emotion prediction result. When the signal-to-noise ratio of the voice data or the occlusion rate of the face image does not meet the preset threshold, the emotion category prediction is performed based on one of the voice data and the face image, which can automatically switch to single-modal emotion prediction, thereby improving the robustness of emotion prediction in noisy environments and face information missing scenarios.

[0078] With reference to Figure 3A In step 303, text data is generated based on the voice data, and subtitle data is generated based on the emotion prediction result and the text data.

[0079] Here, a 4-bit quantized lightweight speech recognition model Whisper-tiny can be loaded, and the speech recognition model is used to transcribe the voice data to text data. The emotion prediction result and the text data are combined to obtain the subtitle data.

[0080] In some embodiments, the subtitle data is generated based on the emotion prediction result and the text data, which can be achieved by the following steps: based on the emotion category corresponding to the emotion prediction result, determining the target display format of the subtitle frame corresponding to the emotion prediction result and the target subtitle color, wherein each emotion category corresponds to a preset display format of the subtitle frame and a preset subtitle color; using the target display format of the subtitle frame and the target subtitle color, rendering the text data to obtain the subtitle data.

[0081] Here, the subtitle data can be displayed in the form of a semi-transparent bubble, and each emotion category corresponds to a preset display format of the subtitle frame and a preset subtitle color. The display format of the subtitle frame includes the subtitle frame color, the subtitle frame line, and the flashing effect of the subtitle frame. The subtitle color is the font color in the subtitle data.

[0082] Based on the emotion category corresponding to the emotion prediction result, the display format of the subtitle border corresponding to the emotion prediction result and the subtitle color, that is, the target display format of the subtitle border and the target subtitle color, are determined. The emotion prediction result, the text data, and the timestamp are combined to obtain triplet data. In response to the data rendering instruction, the triplet data is rendered using the target display format of the subtitle border and the target subtitle color to obtain subtitle data.

[0083] For example, the emotion category corresponding to the emotion prediction result is "anger", and the subtitle data corresponding to the emotion prediction result is displayed as follows: the subtitle border color is red, the subtitle border line is thick, the subtitle border flicker effect is slightly flickering, and the subtitle color is red.

[0084] For example, the emotion category corresponding to the emotion prediction result is "doubt", and the subtitle data corresponding to the emotion prediction result is displayed as follows: the subtitle border color is blue, the subtitle border line is dashed, the subtitle border flicker effect is not flickering, and the subtitle color is blue.

[0085] In the embodiments of the present application, the target display format of the subtitle border corresponding to the emotion prediction result and the target subtitle color are determined, and the text data is rendered using the target display format of the subtitle border and the target subtitle color to obtain subtitle data, which can supplement the emotional information of the subtitle, realize the simultaneous presentation of the content of the subtitle and the fine-grained emotion, improve the richness of the subtitle display, and provide a multi-dimensional understanding channel for hearing aid, cross-language communication, and improve communication efficiency.

[0086] In some embodiments, the following steps can also be performed: when it is detected that the vibration mode is in the starting state, the target vibration frequency corresponding to the terminal is determined based on the emotion category corresponding to the subtitle data, and vibration is performed based on the target vibration frequency, wherein each emotion category corresponds to a preset vibration frequency; when it is detected that the voice reading mode is in the starting state, the subtitle voice corresponding to the subtitle data is output.

[0087] Here, each emotion category corresponds to a preset vibration frequency, and the preset vibration frequency is a frequency at which the terminal is previously set to vibrate. When it is detected that the vibration mode is in the starting state, the frequency at which the terminal vibrates at this time, that is, the target vibration frequency, is determined based on the emotion category corresponding to the subtitle data, so that vibration is performed based on the target vibration frequency. For example, the emotion category corresponding to the subtitle data is "anger", and the preset vibration frequency corresponding to this emotion category is violent vibration. At this time, the target vibration frequency corresponding to the terminal is determined to be violent vibration, and violent vibration is performed based on the target vibration frequency.

[0088] When it is detected that the voice reading mode is in the starting state, the text corresponding to the subtitle data is converted into a subtitle voice, and the subtitle voice is output.

[0089] In this embodiment, based on the emotion category corresponding to the subtitle data, the target vibration frequency corresponding to the terminal is determined, and vibration is performed based on the target vibration frequency. The subtitle voice corresponding to the subtitle data is also output. The subtitle content can be displayed in a visual way through vibration mode and voice reading mode, making it easier for visually impaired or colorblind users to understand the content and tone of the subtitle data.

[0090] Continue to refer to Figure 3A In step 304, eye-tracking data corresponding to the second object and position information of the first object relative to the second object are obtained.

[0091] Here, the second object can be someone who engages in dialogue or communication with the first object, or someone who listens to the first object speak. Eye movement data corresponding to the second object is collected using a built-in eye-tracking sensor. This data includes the eye's center position, eye rotation angle, and fixation time. The direction of the second object is measured using a built-in inertial sensor to obtain the position information of the first object relative to the second object.

[0092] Continue to refer to Figure 3A In step 305, the target position of the subtitle data is determined based on the location information and eye-tracking data, and the subtitle data is displayed at the target position.

[0093] Here, based on location information and eye movement data, the target position of the subtitle data is determined according to the principle of minimum eye saccades, so that the subtitle data can be displayed at the target position. The principle of minimum eye saccades is to minimize the eye movement load of the second object when viewing the subtitle data.

[0094] In some embodiments, eye-tracking data includes the center position of the eyeball, see [link to relevant documentation]. Figure 3C , Figure 3C This is a schematic diagram of the third process of the data processing method provided in the embodiments of this application. Figure 3A In step 305 shown, "determine the target position of the caption data based on location information and eye-tracking data," which can be achieved through... Figure 3C Steps 3051 to 3054 are implemented, and will be explained in detail below.

[0095] In step 3051, the unit vector pointing from the center of the eyeball to the preset gaze point is determined as the gaze vector.

[0096] Here, the preset fixation point is a pre-set point used for gazing at the second object. Vector operations are performed based on the coordinates of the eye center and the coordinates of the preset fixation point to obtain the vector operation result. The vector operation result is then normalized to obtain the unit vector pointing from the eye center to the preset fixation point, which is the gaze vector.

[0097] In step 3052, at least one preset position for displaying the subtitle data is acquired, a first distance between the gaze vector and the preset position is determined for each preset position, and a second distance between the position information and the preset position is determined.

[0098] Here, the at least one preset position for displaying the subtitle data can be determined according to a preference of watching the subtitle data, and the preset position can be a top of a screen, a bottom of the screen, a middle of the screen, a left side of the screen, or a right side of the screen. A difference between the gaze vector and the preset position is determined as the first distance for each preset position. A difference between the position information and the preset position is determined as the second distance. For example, the first distance can be represented as |G-P|, and the second distance can be represented as |S-P|, where G represents the gaze vector, P represents the preset position, and S represents the position information.

[0099] In step 3053, the first distance and the second distance are weighted and summed by using a first preset weight and a second preset weight to obtain a summation result.

[0100] Here, the first preset weight is set to a first value when an angle between the gaze vector and the position information is less than an angle threshold, and the first preset weight is set to a second value when the angle between the gaze vector and the position information is greater than or equal to the angle threshold. The sum of the first preset weight and the second preset weight is 1. The first distance and the second distance are weighted and summed to obtain the summation result.

[0101] For example, the angle threshold is 15°, the first preset weight is set to the first value when the angle between the gaze vector and the position information is less than 15°, the first value is 0.7, the first preset weight is 0.7, and the second preset weight is 0.3. The first preset weight is set to the second value when the angle between the gaze vector and the position information is greater than or equal to 15°, the second value is 0.9, the first preset weight is 0.9, and the second preset weight is 0.1. The summation result can be represented as w1|G-P|+w2|S-P|, where G represents the gaze vector, P represents the preset position, S represents the position information, w1 represents the first preset weight, and w2 represents the second preset weight.

[0102] In step 3054, a preset position corresponding to a minimum summation result is determined as a target position of the subtitle data.

[0103] Here, the preset position that makes the summation result minimum is determined, and the position is determined as the target position of the subtitle data. For example, the target position of the subtitle data can be determined by formula (1):

[0104] P*=argmin(w1|G-P|+w2|S-P|) (1)

[0105] Wherein, P* represents the target position of the subtitle data, G represents the gaze vector, P represents the preset position, S represents the position information, w1 represents the first preset weight, and w2 represents the second preset weight.

[0106] In the embodiments of the present application, the target position of the subtitle data is determined comprehensively in combination with the first distance between the gaze vector and the preset position for displaying the subtitle data, the second distance between the position information of the first object relative to the second object and the preset position, so as to realize adaptive adjustment of the subtitle position, reduce the visual migration load, and thus improve the subtitle reading experience in the multi-person dialogue scene.

[0107] In some embodiments, when the position information of the first object relative to the second object changes, the target position of the subtitle data is moved, which can be realized by the following steps: when the position information of the first object relative to the second object changes, the updated position information is determined; based on the updated position information and the eye movement data, the updated target position of the subtitle data is determined; when the deviation angle between the target position of the subtitle data and the updated target position of the subtitle data is greater than a preset angle, the third distance between the target position of the subtitle data and the updated target position of the subtitle data is determined; based on the third distance and a preset adjustment coefficient, the target position of the subtitle data is moved to the updated target position of the subtitle data.

[0108] Here, when the position information of the first object relative to the second object changes, the direction in which the second object is located is re-measured by the built-in inertial sensor to obtain the updated position information of the first object relative to the second object.

[0109] The unit vector of the eye center position in the eye movement data pointing to the preset fixation point is determined as the gaze vector, the first distance between the gaze vector and the preset position for displaying the subtitle data is determined, and the third distance between the updated position information and the preset position is determined. The first distance and the third distance are weighted and summed by using the first preset weight and the third preset weight to obtain a summation result. The preset position corresponding to the minimum summation result is determined as the updated target position of the subtitle data.

[0110] When the deviation angle between the target position of the subtitle data and the updated target position of the subtitle data is greater than a preset angle, the difference between the target position of the subtitle data and the updated target position of the subtitle data is determined as the third distance. The third distance is multiplied by a preset adjustment coefficient to obtain a position adjustment amount, and the position adjustment amount is summed with the target position of the subtitle data to obtain a migration position of the subtitle. The migration position of the subtitle is used to gradually move the target position of the subtitle data to the updated target position of the subtitle data. In an example, the migration position of the subtitle can be determined by formula (2):

[0111] P_t = P_old + a (P_new - P_old) (2)

[0112] wherein, P_t represents the migration position of the subtitle, P_old represents the target position of the subtitle data, P_new represents the updated target position of the subtitle data, and a represents a preset adjustment coefficient.

[0113] In the embodiments of the present application, the target position of the subtitle data is moved to the updated target position of the subtitle data based on the distance between the target position of the subtitle data and the updated target position of the subtitle data and the preset adjustment coefficient, the gradual migration of the subtitle position is realized, and the influence of the instant jump of the subtitle on the reading experience of the subtitle is avoided.

[0114] In some embodiments, the following steps can also be performed: in response to an adjustment instruction for the font of the subtitle, adjusting the font size of the subtitle data to obtain adjusted subtitle data; or in response to a locking instruction for the position of the subtitle, fixing the target position of the subtitle data and displaying the subtitle data at the target position.

[0115] Here, in response to the adjustment instruction for the font of the subtitle, the adjustment instruction can be a two-finger zoom touch operation applied by the second object on the screen, the font size in the subtitle data is adjusted to obtain adjusted subtitle data. In response to the locking instruction for the position of the subtitle, the locking instruction can be a "clenched fist unclench" gesture operation applied by the second object in front of the screen, the target position of the subtitle data is fixed so that the target position of the subtitle data remains unchanged, and the subtitle data is displayed at the fixed target position.

[0116] In the embodiments of the present application, in response to the adjustment instruction for the font of the subtitle, the font size of the subtitle data is adjusted, and in response to the locking instruction for the position of the subtitle, the target position of the subtitle data is fixed, which can control the font and position of the subtitle in real time, and significantly improve the readability, aesthetic appearance and reading experience of the subtitle.

[0117] In some embodiments, in the process of training the emotion recognition model, the emotion recognition model is trained using local voice data and facial images, the emotion recognition model is fine-tuned locally using Low-Rank Adaptation (LoRA), and the gradient of the emotion recognition model is calculated. The gradient with added Gaussian noise is uploaded to the cloud, the cloud performs noise clipping and aggregation on the gradient with added Gaussian noise to obtain an optimized emotion recognition model through a differentially private federated learning emotion recognition model, and the optimized emotion recognition model is distributed to the local. Differential privacy is a mathematical framework for ensuring that sensitive information about individuals is not disclosed during data analysis. The core idea is to add appropriate noise to the data. Federated learning is a distributed machine learning technique, and differentially private federated learning is a technique that combines differential privacy and federated learning, aiming to protect data privacy while achieving efficient distributed model training.

[0118] In some embodiments, the data processing method provided by the embodiments of the present application can be applied in the field of game live streaming. The voice data and facial images corresponding to the first object are obtained on the game live streaming platform. When it is determined that the emotion prediction condition is met based on the voice data and the facial images, the emotion category prediction is performed based on at least one of the voice data and the facial images to obtain an emotion prediction result. The emotion category of the subtitle can be predicted by combining the voice data and the facial images, the fine-grained emotion of the subtitle is realized, and the accuracy of the emotion prediction result is improved. Then, the text data is generated based on the voice data, and the subtitle data is generated based on the emotion prediction result and the text data. The subtitle can present fine-grained emotion. The eye movement data corresponding to the second object and the position information of the first object relative to the second object are further obtained, and then the target position of the subtitle data is determined based on the position information and the eye movement data, and the subtitle data is displayed at the target position. In this way, the eye movement tracking information and the speaker positioning information are combined to realize adaptive adjustment of the subtitle position, reduce the visual migration load, and thus improve the subtitle reading experience in the multi-player game scenario.

[0119] In the following, the data processing method provided by the embodiments of the present application will be described in the example application in the scene of affective adaptive real-time subtitle generation and layout.

[0120] In the related art, the subtitle can only prompt "laughter, applause" with an emoji, and cannot distinguish fine-grained emotions such as "anger, satire, hesitation" in the subtitle, lacks fine-grained emotion discrimination of the subtitle, and affects the understanding of the subtitle by hearing-impaired or foreign language learners. Displaying the subtitle at a fixed position in the screen, such as displaying the subtitle at a fixed position at the bottom of the screen or a fixed position in the center of the screen, cannot realize adaptive adjustment of the position of the subtitle, resulting in the need for constant eye jumping between the speaking object and the subtitle when multiple people are talking or the user is looking down, which easily causes loss of subtitle information and visual fatigue. Only using the speech features or facial information of the speaking object to recognize the emotion category of the subtitle significantly increases the error rate of emotion discrimination in noisy environments, long-distance sound pickup, and scenes where facial information is missing, and lacks a backup path for emotion reasoning.

[0121] Traditional augmented reality (AR) subtitle glasses upload audio to the cloud to automatically transcribe the audio into subtitles in the cloud, with an average round-trip delay of 400-800 milliseconds, and the call content and facial images are publicly disclosed on the server side, which raises data compliance risks. In addition, the subtitle is misaligned with the position of the speaking object, making it difficult to determine the speaking object corresponding to the subtitle, reducing the practicality of the subtitle.

[0122] Embodiments of the present application address the problems in the related art and propose a data processing method, which includes the following improvements compared to the related art:

[0123] Cross-modal fusion based on speech prosody and facial expression in a mobile terminal accurately identifies the emotion category of the subtitle, automatically degrades to single-modal emotion reasoning in noisy or occluded environments to supplement the emotion information of the subtitle, and can simultaneously present the content and tone of the subtitle. Combined with eye tracking and positioning of the speaking object (the first object in the above embodiment), the position of the subtitle and the speaking object is accurately bound, reducing eye migration and improving the reading experience of the subtitle in a multi-person conversation scenario. The emotion recognition model is optimized using differential privacy federated learning, balancing data compliance and real-time performance, and realizing the collaboration of multiple technical links such as multi-modal emotion recognition, eye-driven dynamic layout of subtitles, and low-latency reasoning on the mobile terminal side.

[0124] On the product level, based on terminals such as smartphones, tablets, computer video conferencing, AR subtitle glasses with forward-facing cameras and eye movement sensors, subtitles are overlaid in real time near the user's (the second object in the above embodiment) line of sight, and the emotions of the speaking object are presented synchronously in multiple channels such as color, icon, and vibration. The position of the subtitle is adaptively adjusted according to the user's gaze point and the orientation of the speaking object, significantly reducing the eye migration load of hearing-impaired people and in noisy environments.

[0125] At the algorithm level, the mobile terminal side extracts the prosodic features (fundamental frequency, energy, speech rate) and facial expression action units simultaneously, uses the graph attention network to cross-modal fusion, outputs the emotion label and inserts the subtitle stream. When the speech signal-to-noise ratio is low or the face occlusion rate is high, it automatically degrades to single-modal emotion reasoning, ensuring the robustness of emotion recognition. The speech transcription uses a 4-bit quantized lightweight speech recognition model Whisper-tiny, and the emotion recognition model uses Low-Rank Adaptation (LoRA) fine-tuning. The average delay on the mobile terminal is less than or equal to 150 milliseconds, and the power consumption increases by less than 1 watt.

[0126] In the embodiments of the present application, the emotion recognition model is fine-tuned locally through differential privacy federated learning, uses local data to train the emotion recognition model, and calculates the gradient of the emotion recognition model. The gradient with added Gaussian noise is uploaded to the cloud, and the cloud aggregates the noisy gradient to optimize the emotion recognition model and issues the optimized emotion recognition model, thereby balancing dialect adaptation and data compliance.

[0127] In the embodiments of the present application, the emotion label of the subtitle uses a "text-emotion-rendering instruction" triple data structure, which is downward compatible with the Web Video Text Tracks (WebVTT) standard and the Accessible Rich Internet Applications (ARIA) standard, and is easy to integrate into scenarios such as conference software, game live streaming, and barrier-free cinema. The system also opens a Software Development Kit (SDK) for third-party applications to call the real-time subtitle and emotion Application Programming Interface (API), achieving consistency in cross-platform experience.

[0128] Below, taking a smartphone as an example, the human-computer interaction process of real-time subtitle generation and layout is described.

[0129] When a user first launches an application (App) in a smartphone, two-step guide interfaces of "Microphone & Camera Authorization" and "Gaze / Speaker Calibration" are popped up in the App in sequence. In the guide interface of "Microphone & Camera Authorization", the system prompts that audio and video data for subtitle generation and emotion recognition need to be collected, and the audio and video data are processed locally, and the user can select the "Only End-Side Processing" function or the "Allow Federated Learning" function. In the guide interface of "Gaze / Speaker Calibration", the screen appears 3 random points for the user to gaze at to establish the user's gaze vector, and then prompts the user to turn his head to read two phrases to establish the speaker object space model, and the whole process can be completed within 30 seconds.

[0130] A "Real-Time Subtitle" switch is displayed at the bottom center of the App interface, and short pressing the "Real-Time Subtitle" switch can start listening to the speech of the speaker object. The sound pickup level is displayed in the top status bar of the App interface, and if the system detects pure media audio (such as songs played by a music player, audio played by a short video App, that is, non-microphone collected speaker object voice), the gradient of the audio feature or the emotion recognition model is not uploaded, and the emotion recognition model is not trained with non-conversation audio. Long pressing the "Real-Time Subtitle" switch can enter the voice transcription settings of multiple languages, supporting automatic detection or manual selection of bilingual subtitles, and the menu of the App interface can be pulled down to switch to a "Only Emotion Labeling" lightweight mode (in this mode, voice transcription is still performed, but the transcribed text is not rendered, only emotion labeling metadata is retained, and the emotion is labeled with color / small icons at the speaker object position, and the transcribed text is discarded immediately) to adapt to low battery scenarios.

[0131] When the speaker object is detected, the subtitle appears in the form of a semi-transparent bubble on the side of the speaker object's shoulder, and the color and border style of the subtitle bubble depend on the emotion label, such as: the subtitle bubble corresponding to the emotion label "Anger" is displayed as a red thick border and slightly flickers, the subtitle bubble corresponding to the emotion label "Question" is displayed as a blue dotted border, and the subtitle bubble corresponding to the emotion label "Neutral" is displayed as a gray-white border. If the user's gaze point deviates from the speaker object position by more than a preset angle, the system smoothly moves the subtitle bubble at a speed of 200 milliseconds, ensuring that the subtitle is always within the user's 3-5 degree field of view. If multiple speaker objects speak at the same time, the subtitle bubble is placed on top in time sequence and a subtitle bubble list appears in the lower left corner of the App interface for the user to manually click to lock.

[0132] The user can expand the full-text history (full-text log of the subtitles since the current session, including text, emotion label, and timestamp) by swiping down or double-clicking the subtitles in the App interface, which facilitates reviewing the missed subtitle content. Double-clicking the subtitles again can fold the full-text history. The user can adjust the font size of the subtitles by two-finger zooming in the App interface, and quickly mute the speaker for 3 minutes by swiping left. In the AR glasses application scenario, the user can lock the current subtitle position by the "clenching and unclenching" gesture, avoiding subtitle jumping in a noisy environment. For users with weak vision or color blindness, the subtitles are accompanied by optional vibration prompts, such as higher emotion levels and faster vibration frequencies. In the call mode, voice reading reverse output (reading the subtitle content) can be enabled, which is helpful for the deaf-blind double-impaired to understand the subtitle content.

[0133] The App settings interface provides four switching items: "local mode / federated learning mode", "emotion granularity", "subtitle layout strategy", and "data cleaning". The setting option of "local mode / federated learning mode" is the local mode by default. After enabling the federated learning mode, the system only transmits the gradient information of the emotion recognition model in the network connection state and the charging state. The setting option of "emotion granularity" is basic 5 categories or extended 12 categories. The higher the emotion granularity, the longer the calculation time. The setting option of "subtitle layout strategy" is automatic strategy, fixed bottom bar strategy, and fixed central strategy. In the automatic strategy, the system calculates the subtitle position in real time according to the minimum eye jump formula. "Data cleaning" is a one-key cleaning of local cached audio data and intermediate emotion tensors.

[0134] When the system detects that the voice signal-to-noise ratio of the voice data is low or the face occlusion rate of the face image is high, a yellow "low quality" mark appears in the status bar of the App interface, and automatically switches to single-modal emotion recognition. If there is no valid emotion output for 5 seconds, only the pure text subtitle is retained. The user can manually switch to "pure text mode" to save energy, and the switching button is located in the upper right corner of the App main interface.

[0135] The embodiments of the present application can be applied to online meetings, street conversations, barrier-free cinemas, and other scenarios. In the online meeting scenario, through the video conference plug-in mode, the subtitle bubble is bound to the lower left corner of the video window of each participant. Mouse hovering over the subtitle bubble can display the complete subtitle content. In the real-time conversation scenario, by wearing AR glasses, the subtitles on the AR glasses move with the position of the speaker. If the user bends down to operate, the subtitles automatically move to the lower right corner of the lens and shrink to a single line scrolling. In the barrier-free cinema scenario, the smartphone is placed horizontally on the armrest, and the subtitles move left and right with the character's position. When the emotional intensity is high, a vibration reminder is given.

[0136] The following describes the overall technical architecture provided by the embodiments of the present application. The system adds a real-time subtitle service layer in the mobile terminal, which includes seven functional layers connected in series, namely, a multi-source collection layer, a single-modal feature extraction layer, a cross-modal fusion and emotion reasoning layer, a speech transcription layer, a subtitle synthesis layer, a dynamic layout engine layer, and a rendering and interaction layer.

[0137] In the multi-source collection layer, audio is collected by a microphone array at a frequency of 16 kHz and a frame rate of 20 milliseconds, a face image is captured by a front-facing camera at a rate of 30 frames per second, a gaze vector is updated once every 8 milliseconds by an eye tracking sensor, and inertial sensor (IMU) data is used to locate the position of the speaker. The gaze vector can be determined in various ways. For AR glasses or a front-facing camera of a smartphone with an auxiliary light source, the vector from the camera's optical center to the center of the pupil, i.e., the gaze vector, is calculated. For single-camera devices, a convolutional neural network or a regression model can be used to estimate the gaze vector based on head posture (IMU measurements) and the user's historical interaction positions.

[0138] In the single-modal feature extraction layer, the fundamental frequency, energy, zero-crossing rate, speech rate, spectral features, and 512-dimensional acoustic embedding of the audio data are extracted as prosodic features (audio features in the above embodiments). The 17 Action Unit (AU) intensity features of the video data (facial features in the above embodiments) are extracted by a lightweight face recognition model MobileFace2, and then the timestamps of the two types of features are unified.

[0139] In the cross-modal fusion and emotion reasoning layer, the system aggregates the 17 Action Unit intensity features and 6-dimensional prosodic features in 200-millisecond sliding windows to form 23-dimensional feature nodes in each window. A cross-modal emotion graph is constructed using time adjacency and cross-modal mutual information, and a two-layer structure of a Graph Attention Network (GAT) is used to obtain a 12-dimensional emotion vector. The 12-dimensional emotion vector is normalized using a softmax function to obtain the probability distribution of 12 emotion categories. The 12 emotion categories include joy, calm, surprise, sadness, anger, disgust, fear, doubt, sarcasm, embarrassment, fatigue, and excitement. The emotion category with the maximum probability is taken as the final emotion label, and the emotion category is adaptively mapped to five emotion intensity levels according to the quantile of the maximum probability value. The five emotion intensity levels include extremely negative, negative, neutral, positive, and extremely positive.

[0140] In the speech transcription layer, an incremental text is continuously generated by a lightweight speech recognition model Whisper-tiny with 4-bit quantization.

[0141] In the subtitle synthesis layer, the text, emotion, and timestamp are packaged as a triple, and rendering instructions such as color, border, and vibration are attached.

[0142] In the dynamic layout engine layer, the gaze vector of the user and the direction of the speaking object (the position information of the first object relative to the second object in the above embodiment) are combined to calculate the subtitle anchor point according to the minimum saccade principle. The subtitle anchor point is the floating position coordinate of the subtitle presentation on the device screen or the field of view of the glasses. The minimum saccade principle is to reduce the user's eye movement load and dynamically adjust the subtitle position, with the optimization goal of minimizing the angle of the field of view between the user's gaze vector and the subtitle anchor point, and the subtitle anchor point being as close as possible to the direction of the speaking object. In an example, the target position of the subtitle can be determined by formula (1):

[0143] P*=argmin(w1|G-P|+w2|S-P|) (1)

[0144] where P* represents the target position of the subtitle, G represents the gaze vector, P represents the preset position of the subtitle, represents, S represents the direction of the speaking object, w1 represents the first preset weight, and w2 represents the second preset weight.

[0145] The system obtains the current user gaze vector and the direction of the speaking object every 50 milliseconds. When the angle between the user gaze vector and the direction of the speaking object is less than 15°, the first preset weight w1 is set to 0.7, and the speaking object is preferentially approached. When the angle between the user gaze vector and the direction of the speaking object is greater than 15°, the first preset weight w1 is set to 0.9, and the user gaze direction is preferentially approached. The updated target position of the subtitle is subjected to exponential smoothing processing, and then input to the rendering queue to update the rendering queue.

[0146] If the angle difference between the updated target position of the subtitle and the target position of the subtitle before the update exceeds 6°, a smoothing algorithm is enabled to perform gradual migration. In an example, the migration position of the subtitle can be determined by formula (2):

[0147] P_t=P_old+α(P_new-P_old) (2)

[0148] where P_t represents the migration position of the subtitle, P_old represents the target position of the subtitle before the update, P_new represents the updated target position of the subtitle, and a represents a preset adjustment coefficient.

[0149] Alpha can be 0.15, the subtitle position is updated once per frame, and the subtitle position migration is completed in about 150 milliseconds to avoid the impact of subtitle instantaneous jumping on the user's subtitle reading experience.

[0150] In the rendering and interaction layer, subtitle drawing is completed through a 16-millisecond refresh cycle per frame, and gestures or touch operations are received. Users are allowed to quickly adjust the font and position of the subtitle, shield or highlight the subtitle of a certain speaking object, or lock the current subtitle position in the AR scene. Gestures or touch operations are captured through the front-end interface, and the local state machine updates the subtitle rendering parameters or the local model sets the subtitle rendering parameters, without uploading private data such as audio and video to the server.

[0151] In the embodiments of the present application, the time consumption of feature extraction on audio and video frames is about 2 milliseconds, and the time consumption of unifying the timestamps of the two types of features is 1 millisecond. The emotion reasoning consumes 8 milliseconds, the average time consumption of speech incremental transcription is 40 milliseconds, the time consumption of subtitle synthesis and layout is 3 milliseconds, the time consumption of subtitle rendering is within 16 milliseconds, and the total end-to-end delay time is between 70-150 milliseconds, which varies with the speed of the audio.

[0152] In the embodiments of the present application, emotion recognition is performed by a local graph attention network that fuses speech and facial images to achieve stable emotion discrimination in noisy or occluded environments. The subtitle adopts a "text-emotion-rendering instruction" triple structure, laying the foundation for multi-channel presentation and standard interfaces. The subtitle layout algorithm combines gaze direction and speaker direction as double factors to dynamically adjust the subtitle position, which can reduce the saccade distance by 30%-40%. The entire reasoning process is completed on a mobile terminal, and the emotion recognition model is updated using differential privacy federated learning, which not only meets the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA), but also solves the 400-800 millisecond cloud delay.

[0153] In the embodiments of the present application, the network topology calculation mode includes a full end-side mode, an end-cloud hybrid mode, and a near-end edge mode. The full end-side mode has a total size of about 30 MB for all models, which runs on a Neural Network Processing Unit (NPU) of a smartphone or glasses, and is suitable for privacy-sensitive scenarios. The end-cloud hybrid mode performs emotion reasoning on the end side, and uploads the generated subtitle triplets to a conference server, with a bandwidth requirement of 2 kilobytes per second. The near-end edge mode deploys a lightweight speech recognition model to an edge computing node in a network connection state, and the smartphone only performs emotion reasoning and subtitle layout processes, with a time delay further reduced to 60 milliseconds.

[0154] In the embodiments of the present application, when the voice signal-to-noise ratio is lower than 10 decibels, the system reduces the prosody feature weight and performs single-modal emotion recognition on the face image. When the face occlusion rate of the face image exceeds 40%, single-modal emotion recognition is performed on the audio. If the voice signal-to-noise ratio is lower than 10 decibels and the face occlusion rate of the face image exceeds 40%, only the pure text subtitle is output. When the NPU occupancy rate is higher than 80%, the system automatically switches to an 8-bit quantization emotion recognition model, and the emotion categories are reduced from 12 to 5, so as to save system resources.

[0155] In the embodiments of the present application, the system defaults that the local emotion recognition model is fine-tuned by Low-Rank Adaptation (LoRA) to adapt to dialects. The gradient information of the emotion recognition model is uploaded after being added with noise by differential privacy, and the server aggregates the gradient information every week and delivers the optimized emotion recognition model. The user can start the "zero upload" mode at any time to prohibit the audio and video data from being exported.

[0156] In the embodiments of the present application, the emotion recognition model can use a multi-modal self-attention mechanism network Transformer architecture or a dual-tower architecture to replace the graph attention network. The single-camera device that locates the direction of the speaking object and the inertial sensor can be upgraded to a dual-camera stereo ranging or Ultra-Wideband (UWB) tag. The eye movement tracking can be performed by using a screen-under Time of Flight (ToF) or an inertial sensor. The subtitle rendering layer can be switched to a Web Graphics Library (WebGL) or a cross-platform development tool (Unity) overlay layer. The differential privacy can be replaced by a secure execution environment or homomorphic encryption aggregation, and the end-side inference hardware can also use a Bluetooth computing stick or a Field Programmable Gate Array (FPGA) acceleration card. Through the above modularization and multi-path replacement design, the system can not only be applied to mobile terminals such as smartphones and AR glasses, but also has expandable space for algorithms, hardware and networks.

[0157] In the embodiments of the application, voice transcription can be performed first, and then emotion reasoning can be performed. The binding of text and emotion during subtitle generation can reduce reasoning waiting time and is suitable for ultra-high-speed speech scenarios. Multi-modal streaming media data (such as heart rate photoplethysmography information and inertial gait) can be used as auxiliary signals for emotion reasoning to improve the robustness of emotion recognition. Depth camera or millimeter wave point cloud can be used to estimate facial expression action units, and a user gaze vector can be inferred through an inertial measurement model, which is suitable for low-end devices without a dedicated eye movement sensor. The subtitle anchoring logic can be adjusted according to the cognitive load or the priority of the user-defined subtitle position (such as fixing the subtitle position at the top left corner of the interface). The embodiments of the application can also be applied to e-sports live streaming, online education, court records, and automatically generated movie subtitles, etc. to label the emotional information of the subtitles.

[0158] In the above-mentioned emotion-adaptive real-time subtitle generation and layout scenarios, through cross-modal fusion of speech prosody and expression action units, the system can still output 12 categories and 5 levels of intensity emotion labels under noisy or partially occluded conditions. The emotion recognition accuracy of the end-side emotion reasoning is improved by 9.4 percentage points compared to the convolutional neural network model. The subtitle understanding accuracy of the hearing-impaired is increased from 67% to 92%, and the complete presentation of the emotional context is realized. The dynamic layout engine adjusts the subtitle position in real time according to the minimum eye jump principle. In multi-person dialogue scenarios, the average eye jump distance is reduced from 14.7° to 9.8°, the eye jump amplitude is reduced by 33%, the user subjective fatigue score is reduced by two levels, and the visual migration load is significantly reduced. The end-to-end reasoning delay time is 70-150 milliseconds, which is 4-6 times shorter than the round-trip time delay of the cloud-based transcription solution. The reasoning power consumption of the 4-bit quantized model is only increased by 0.9 watts. Smartphones or AR glasses can be used continuously for more than 4 hours, realizing low latency and low power consumption in an end-to-end manner.

[0159] When high noise or high occlusion is detected, the system automatically switches to single-modal emotion reasoning or reverts to outputting pure text subtitles. The subtitle availability remains above 99%. When the NPU load is too high, the emotion categories are automatically reduced in dimension, realizing a strong robustness degradation strategy. Federated learning combined with differential privacy noise clipping reduces the size of the model gradient uploaded at a time to less than 50 kilobytes, balancing dialect adaptation and data compliance. In medical and judicial scenarios, the "zero upload" mode can be enabled to ensure that personal data runs locally in a closed environment, realizing data privacy and regulatory compliance. The subtitle triple (text-emotion-rendering instruction) is downward compatible with WebVTT and ARIA standards, which can be seamlessly embedded in various applications such as conference software, game live streaming, cinema projection, etc. The SDK size is less than 5 megabytes, improving the secondary development efficiency and realizing easy integration and cross-scene compatibility of subtitles. Through emotion perception, gaze-adaptive layout, and local privacy computing, the completeness, comfort, real-time performance, and compliance of the subtitle experience are comprehensively improved, providing practical and quantifiable technical value for hearing-impaired assistance, cross-language communication, and communication in noisy environments.

[0160] The following continues to illustrate an exemplary structure of the data processing apparatus 455 provided by the embodiments of the present application, which is implemented as a software module, in some embodiments, as shown in the figure, the software module stored in the data processing apparatus 455 of the memory 450 can include: a first acquisition module 4551 configured to acquire voice data and facial images corresponding to a first object; an emotion prediction module 4552 configured to, when it is determined that an emotion prediction condition is met based on the voice data and the facial images, perform emotion category prediction based on at least one of the voice data and the facial images to obtain an emotion prediction result; a data generation module 4553 configured to generate text data based on the voice data, and generate subtitle data based on the emotion prediction result and the text data; a second acquisition module 4554 configured to acquire eye movement data corresponding to a second object, and position information of the first object relative to the second object; and a position determination module 4555 configured to determine a target position of the subtitle data based on the position information and the eye movement data, and display the subtitle data at the target position. Figure 2

[0161] In some embodiments, the emotion prediction module 4552 is further configured to determine a signal-to-noise ratio of the voice data and an occlusion rate of the facial images; and determine that the emotion prediction condition is met when the signal-to-noise ratio of the voice data is greater than or equal to a first preset threshold, or the occlusion rate of the facial images is less than or equal to a second preset threshold.

[0162] In some embodiments, the emotion prediction module 4552 is further configured to, when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the facial images is less than or equal to the second preset threshold, perform emotion category prediction based on the voice data and the facial images to obtain the emotion prediction result; when the signal-to-noise ratio of the voice data is less than the first preset threshold, and the occlusion rate of the facial images is less than or equal to the second preset threshold, perform emotion category prediction based on the facial images to obtain the emotion prediction result; and when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the facial images is greater than the second preset threshold, perform emotion category prediction based on the voice data to obtain the emotion prediction result.

[0163] In some embodiments, the emotion prediction module 4552 is further configured to perform feature extraction on the voice data and the facial images respectively to obtain audio features and facial features respectively; fuse the audio features and the facial features to obtain fused features; perform emotion category prediction on the fused features by using a graph attention network to obtain an emotion category probability distribution; and perform emotion intensity mapping based on the emotion category probability distribution to obtain the emotion prediction result.

[0164] ​In some embodiments, the data generation module 4553 is further configured to determine a target display format of a subtitle frame and a target subtitle color corresponding to the emotion prediction result based on an emotion category corresponding to the emotion prediction result, wherein each emotion category corresponds to a preset display format of the subtitle frame and a preset subtitle color; and render the text data to obtain the subtitle data by using the target display format of the subtitle frame and the target subtitle color.

[0165] In some embodiments, the data generation module 4553 is further configured to determine a target vibration frequency of the terminal based on an emotion category corresponding to the subtitle data when it is detected that the vibration mode is in an enabled state, and vibrate based on the target vibration frequency, wherein each emotion category corresponds to a preset vibration frequency; and output a subtitle voice corresponding to the subtitle data when it is detected that the voice reading mode is in the enabled state.

[0166] In some embodiments, the eye movement data includes a center position of an eyeball, and the position determination module 4555 is further configured to determine a unit vector of the center position of the eyeball pointing to a preset fixation point as a gaze vector; obtain at least one preset position for displaying the subtitle data, determine a first distance between the gaze vector and each preset position and a second distance between the position information and each preset position for each preset position; obtain a summation result by using a first preset weight and a second preset weight to perform weighted summation on the first distance and the second distance; and determine a preset position corresponding to a minimum summation result as a target position of the subtitle data.

[0167] In some embodiments, the position determination module 4555 is further configured to determine updated position information when the position information of the first object relative to the second object changes; determine an updated target position of the subtitle data based on the updated position information and the eye movement data; determine a third distance between the target position of the subtitle data and the updated target position of the subtitle data when an angle of deviation between the target position of the subtitle data and the updated target position of the subtitle data is greater than a preset angle; and move the target position of the subtitle data to the updated target position of the subtitle data based on the third distance and a preset adjustment coefficient.

[0168] In some embodiments, the position determination module 4555 is further configured to perform font size adjustment on the subtitle data to obtain adjusted subtitle data in response to an adjustment instruction for a subtitle font; or fix the target position of the subtitle data and display the subtitle data at the target position in response to a locking instruction for a subtitle position.

[0169] The embodiment of the present application provides a computer program product, which comprises computer executable instructions or computer programs stored in a computer readable storage medium. The processor of an electronic device reads the computer executable instructions or computer programs from the computer readable storage medium, and the processor executes the computer executable instructions or computer programs, so that the electronic device executes the data processing method provided by the embodiment of the present application.

[0170] The embodiment of the present application provides a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions or computer programs are stored, and when the computer executable instructions or computer programs are executed by a processor, the processor will execute the data processing method provided by the embodiment of the present application, for example, the data processing method shown in the figure. Figure 3A

[0171] In some embodiments, the computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM and the like memory; and can also be various devices including one or any combination of the above memories.

[0172] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.

[0173] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in a part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperative files (for example, files storing one or more modules, subroutines or code parts).

[0174] As an example, the computer executable instructions can be deployed to execute on one electronic device, or on multiple electronic devices located in one place, or on multiple electronic devices distributed in multiple places and interconnected through a communication network.

[0175] ​The above merely provides an example of the present application, but is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, and improvement made within the spirit and scope of the present application shall be included in the protection scope of the present application.

Claims

1. A data processing method, characterized by, The method comprises: obtaining voice data and facial images corresponding to a first object; when it is determined that an emotion prediction condition is met based on the voice data and the facial images, performing emotion category prediction based on at least one of the voice data and the facial images to obtain an emotion prediction result; generating text data based on the voice data, and generating subtitle data based on the emotion prediction result and the text data; obtaining eye movement data corresponding to a second object, and position information of the first object relative to the second object; based on the position information and the eye movement data, determining a target position of the subtitle data, and displaying the subtitle data at the target position.

2. The method of claim 1, wherein, The method further comprises: determining a signal-to-noise ratio of the voice data and an occlusion rate of the facial images; when the signal-to-noise ratio of the voice data is greater than or equal to a first preset threshold, or the occlusion rate of the facial images is less than or equal to a second preset threshold, it is determined that the emotion prediction condition is met.

3. The method of claim 2, wherein, When it is determined that an emotion prediction condition is met based on the voice data and the facial images, performing emotion category prediction based on at least one of the voice data and the facial images to obtain an emotion prediction result, comprising: when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the facial images is less than or equal to the second preset threshold, performing emotion category prediction based on the voice data and the facial images to obtain an emotion prediction result; when the signal-to-noise ratio of the voice data is less than the first preset threshold, and the occlusion rate of the facial images is less than or equal to the second preset threshold, performing emotion category prediction based on the facial images to obtain an emotion prediction result; when the signal-to-noise ratio of the voice data is greater than or equal to the first preset threshold, and the occlusion rate of the facial images is greater than the second preset threshold, performing emotion category prediction based on the voice data to obtain an emotion prediction result.

4. The method of claim 3, wherein, The emotion category prediction based on the voice data and the facial images to obtain an emotion prediction result comprises: performing feature extraction on the voice data and the facial images respectively to obtain audio features and facial features respectively; fusing the audio features and the facial features to obtain fused features; using a graph attention network to perform emotion category prediction on the fused features to obtain an emotion category probability distribution; performing emotion intensity mapping based on the emotion category probability distribution to obtain an emotion prediction result.

5. The method of claim 1, wherein, The generation of subtitle data based on the emotion prediction result and the text data comprises: determining a target display format of a subtitle frame corresponding to the emotion prediction result and a target subtitle color based on an emotion category corresponding to the emotion prediction result, wherein each emotion category corresponds to a preset display format of a subtitle frame and a preset subtitle color; using the target display format of the subtitle frame and the target subtitle color to render the text data to obtain subtitle data.

6. The method of claim 5, wherein, The method further comprises: When it is detected that the vibration mode is in the starting state, a target vibration frequency of the terminal is determined based on an emotion category corresponding to the subtitle data, and the terminal vibrates based on the target vibration frequency, wherein each emotion category corresponds to a preset vibration frequency; When it is detected that the voice reading mode is in the starting state, a subtitle voice corresponding to the subtitle data is output.

7. The method of claim 1, wherein, The eye movement data includes a center position of an eyeball, and the target position of the subtitle data is determined based on the position information and the eye movement data, including: A unit vector of the center position of the eyeball pointing to a preset fixation point is determined as a gaze vector; At least one preset position for displaying the subtitle data is obtained, and for each preset position, a first distance between the gaze vector and the preset position and a second distance between the position information and the preset position are determined; The first distance and the second distance are weighted and summed by using a first preset weight and a second preset weight to obtain a summation result; A preset position corresponding to a minimum summation result is determined as the target position of the subtitle data.

8. The method of claim 7, wherein, The method further includes: When the position information of the first object relative to the second object changes, updated position information is determined; The updated target position of the subtitle data is determined based on the updated position information and the eye movement data; When an angle of deviation between the target position of the subtitle data and the updated target position of the subtitle data is greater than a preset angle, a third distance between the target position of the subtitle data and the updated target position of the subtitle data is determined; The target position of the subtitle data is moved to the updated target position of the subtitle data based on the third distance and a preset adjustment coefficient.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: In response to an adjustment instruction for a subtitle font, the font size of the subtitle data is adjusted to obtain adjusted subtitle data; or In response to a locking instruction for a subtitle position, the target position of the subtitle data is fixed, and the subtitle data is displayed at the target position.

10. A data processing apparatus, characterized by The device includes: A first obtaining module is configured to obtain voice data and a facial image corresponding to a first object; An emotion prediction module is configured to, when it is determined that an emotion prediction condition is met based on the voice data and the facial image, perform emotion category prediction based on at least one of the voice data and the facial image to obtain an emotion prediction result; A data generation module is configured to generate text data based on the voice data and generate subtitle data based on the emotion prediction result and the text data; A second obtaining module is configured to obtain eye movement data corresponding to a second object and position information of the first object relative to the second object; A position determination module is configured to determine a target position of the subtitle data based on the position information and the eye movement data and display the subtitle data at the target position.