Video picture generation method and device, equipment, medium and product

By generating personalized video images to assist users in reading, the problems of constrained imagination space and subjective understanding differences in e-reading are solved, and the reading experience is improved.

CN120475231APending Publication Date: 2025-08-12MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510548078.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The lack of visual assistance in existing electronic reading methods leads to problems such as limited imagination space, subjective understanding differences and obscure reading experience when reading works in grand scenes or complex plots.

Method used

By obtaining the user's current environmental factors, reading historical data, current emotional data and reading area information, a personalized video picture is generated using the multimodal fusion model to assist users in reading.

Benefits of technology

It improves the user's reading experience, makes the reading process more vivid and interesting, and solves the problems of constrained imagination space and differences in subjective understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475231A_ABST
    Figure CN120475231A_ABST
Patent Text Reader

Abstract

The invention discloses a video picture generation method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a current environment factor and reading historical data of a user; identifying current emotion data of the user; positioning a current reading area of the user to obtain text information of the current reading area; and generating a video picture according to the text information, the current emotion data, the current environment factor and reading historical data. A personalized video picture can be generated according to current reading text information, current emotion data, current environment factors and reading historical data of a user, the user is assisted in reading through video vision, the problems of limited imagination space, subjective understanding difference, obscure reading experience and the like in the reading process of the user are solved, reading is more vivid and interesting, and the user experience is improved. And user reading experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic reading technology, and in particular to a method, device, equipment, medium and product for generating a video image. Background Art

[0002] With the advancement of technology, more and more users are reading digital e-books through e-readers, tablets, smartphones, or web browsers. Users select a specific book and use a program to simulate the reading experience of a physical book, presenting the book's contents on the screen. Interactive features for reading digital e-books primarily include page turning with gestures and swiping up and down. During reading, users can browse other users' comments on the content and leave their own comments. E-readers also offer advanced features such as bookmark management, reading progress tracking, and font style adjustment to meet users' personalized reading needs. Users can choose different fonts, sizes, and page layouts to optimize their reading experience. Some e-readers also offer text-to-speech functions, allowing users to listen to the book's contents. While existing e-reading methods offer a relatively user-friendly reading experience, they lack visual support. Users relying solely on textual descriptions to understand and visualize the plot may not fully grasp the author's grand vision. This is especially true when reading works with grand scenes or complex plots, which can be difficult or difficult to understand. Furthermore, each user's understanding and imagination of textual descriptions varies depending on their personal experience, cultural background, and emotional state. Therefore, traditional reading scenarios lack visual assistance, and users rely solely on text descriptions to understand and imagine the story, which can easily lead to problems such as limited imagination, differences in subjective understanding, and obscure reading experience. Summary of the Invention

[0003] The present invention provides a method, device, equipment, medium and product for generating video images, which generate video images based on the user's current reading text information, current emotional data, current environmental factors and reading history data, and assist the user in reading through video vision, thereby solving problems such as limited imagination space, subjective understanding differences and obscure reading experience during the user's reading process.

[0004] To achieve the above objectives, in a first aspect, an embodiment of the present invention provides a method for generating a video image, comprising:

[0005] Obtain the user's current environment and reading history data;

[0006] identifying current emotion data of the user;

[0007] Locating the current reading area of the user and obtaining text information of the current reading area;

[0008] A video image is generated according to the text information, the current emotion data, the current environmental factors and the reading history data.

[0009] As an improvement to the above solution, locating the current reading area of the user and obtaining text information of the current reading area includes:

[0010] Determining a first reading area of the user according to current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user;

[0011] determining a second reading area of the user according to current gesture touch data of the user on the user terminal;

[0012] determining a current reading area of the user according to the first reading area and the second reading area;

[0013] The content of the current reading area is converted into text information.

[0014] As an improvement to the above solution, determining the first reading area of the user based on the current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user includes:

[0015] A gyroscope sensor is used to obtain the current screen posture information of the user terminal where the reading interface is located;

[0016] Determining the user's current line of sight using eye tracking technology, and locating the user's gaze position based on the current line of sight;

[0017] A first reading area of the user is determined according to the current screen posture information and the gaze position.

[0018] As an improvement to the above solution, determining the second reading area of the user based on the current gesture touch data of the user on the user terminal includes:

[0019] Acquiring current gesture touch data of the user on the user terminal;

[0020] Analyzing the current gesture touch data using gesture operation recognition technology to obtain the user's gesture operation and a movement trajectory of the gesture operation on the screen of the user terminal;

[0021] A second reading area of the user is determined according to the gesture operation and the movement trajectory.

[0022] As an improvement to the above solution, the step of generating a video image based on the text information, the current emotion data, the current environmental factors, and the reading history data includes:

[0023] The text information, the current emotion data, the current environmental factors and the reading history data are used to generate a video image and input it into the trained multimodal fusion model;

[0024] generating video content according to the text information;

[0025] generating a presentation style of the video content according to the current emotion data, the current environmental factors, and the reading history data;

[0026] The video content and the presentation style generate a video picture.

[0027] As an improvement to the above solution, the identifying the current emotion data of the user includes:

[0028] Obtaining the current facial information of the user;

[0029] extracting facial expression features based on the current facial information;

[0030] The user's current emotion data is obtained according to the facial expression features.

[0031] In order to achieve the above-mentioned objective, in a second aspect, an embodiment of the present invention provides a video image generating device, comprising:

[0032] Related data acquisition module, used to obtain the user's current environment factors and reading history data;

[0033] An emotion data recognition module, configured to recognize the user's current emotion data;

[0034] A text information acquisition module, configured to locate the current reading area of the user and obtain text information of the current reading area;

[0035] The video frame generation module is used to generate a video frame according to the text information, the current emotional data, the current environmental factors and the reading history data.

[0036] In order to achieve the above-mentioned purpose, in the third aspect, an embodiment of the present invention provides a video picture generating device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the above-mentioned video picture generating method when executing the computer program.

[0037] In order to achieve the above-mentioned purpose, in a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned video screen generation method.

[0038] In addition, to achieve the above-mentioned purpose, an embodiment of the present invention further provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the above-mentioned video image generation method.

[0039] Compared to the prior art, the embodiments of the present invention disclose a method, device, equipment, medium, and product for generating a video image. These methods obtain the user's current environmental factors and reading history data; identify the user's current emotional data; locate the user's current reading area and obtain text information in the current reading area; and generate a video image based on the text information, the current emotional data, the current environmental factors, and the reading history data. This method can generate personalized video images based on the user's current reading text information, current emotional data, current environmental factors, and reading history data, and visually assist the user in reading through video, resolving issues such as limited imagination, subjective understanding differences, and an obscure reading experience during the user's reading process. This makes reading more vivid and interesting, enhancing the user's reading experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 1 is a flow chart of a method for generating a video image provided by an embodiment of the present invention;

[0041] Figure 2 This is a block diagram of a video image generation process provided by an embodiment of the present invention;

[0042] Figure 3 This is a flowchart of a process for obtaining user reading content provided by an embodiment of the present invention;

[0043] Figure 4 This is another flowchart of a user reading content acquisition process provided by an embodiment of the present invention;

[0044] Figure 5 This is a flowchart of a user emotion recognition process provided by an embodiment of the present invention;

[0045] Figure 6 This is a flowchart of an environmental sound feature recognition process provided by an embodiment of the present invention;

[0046] Figure 7 This is another block diagram of a video image generation process provided by an embodiment of the present invention;

[0047] Figure 8 This is a structural diagram of a video image generating device provided by an embodiment of the present invention;

[0048] Figure 9 This is a structural block diagram of a video image generating device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] It should be noted that the terms "comprises" and "specifically" and any variations thereof in the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or apparatuses.

[0051] See also Figure 1 , Figure 1 1 is a flow chart of a method for generating a video image provided by an embodiment of the present invention, the method comprising:

[0052] S1, obtain the user's current environment factors and reading history data;

[0053] S2, identifying the user's current emotion data;

[0054] S3, locating the current reading area of the user and obtaining text information of the current reading area;

[0055] S4, generating a video image based on the text information, the current emotion data, the current environmental factors and the reading history data.

[0056] Exemplarily, the video screen generation method is implemented by a video screen generation server. The video screen generation method described in the embodiment of the present invention can be applied to electronic reading of novels, comics, essays, etc. The video screen generation server can be installed in an electronic reading server (such as an e-reader, tablet computer, smart phone or computer web page, etc.), or exist independently. The video screen generation server can interact with users and interact with electronic reading servers.

[0057] Specifically, the current environmental factors include ambient sound (quiet, noisy, etc.), current time (night, daytime, afternoon tea time, etc.), geographical location (subway, office building, home, mountain, etc.), weather (sunny, rainy, etc.), light, etc.

[0058] The reading history data includes user browsing history records and user interaction behavior data (such as likes, favorites, comments, etc.) and reading style preferences (such as font, font size and page layout settings).

[0059] Furthermore, after generating a video image based on the text information, the current emotion data, the current environmental factors and the reading history data, the method further includes:

[0060] The video image is synchronously presented above (or below) the current reading area.

[0061] In a specific implementation, Figure 2 As shown, Figure 2 This is a flowchart of a video image generation process provided by an embodiment of the present invention. After a user enters the reading page of an electronic reading server, he or she can browse the contents of a book by turning pages or swiping up through gestures. The video image generation server first obtains the text content currently being browsed by the user and obtains the text information on the screen by identifying the position of the user's eyeballs. Secondly, it analyzes the user's current emotions through the user's facial expressions, obtains the current time and location in real time, and captures the current ambient sound information through a microphone. In addition, it also collects and analyzes the user's historical browsing records and user interaction behavior data, such as likes, favorites, and comments, to further understand the user's interests and preferences. Based on this information, the video image generation server can more accurately determine the content that the user is interested in, thereby generating relevant video images. The video image corresponding to the text will automatically appear on the screen of the electronic reading server, or the user can turn on the video assist mode, and the video image will be synchronously presented on the screen of the electronic reading server to enhance the user's reading experience.

[0062] It is worth noting that the following factors may affect the user's reading willingness: (1) Video content quality: The quality of video content is the key to attracting users to continue reading; high-quality videos include clear pictures and attractive content, which can improve users' reading experience and satisfaction; (2) Video experience: The video playback experience is crucial for users to continue reading; fast loading, smooth playback, and no advertising interference can improve users' reading experience and increase user retention rate; (2) Personalized video recommendation: Based on users' reading preferences and historical behaviors, personalized video recommendations can be provided to better meet users' reading needs and increase users' continuous reading time; (4) Video editing tools: Provide simple and easy-to-use video editing tools so that users can edit and produce video content by themselves, increase user participation and creativity, and thus enhance user stickiness; (5) Video sharing and social interaction: Users can enhance their sense of participation and social relationships by sharing videos and participating in video-related social interactions, such as commenting and liking, and promote continuous reading.

[0063] Specifically, step S3 includes:

[0064] S31, determining a first reading area of the user according to current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user;

[0065] S32, determining a second reading area of the user according to current touch gesture data of the user on the user terminal;

[0066] S33, determining a current reading area of the user according to the first reading area and the second reading area;

[0067] S34, converting the content of the current reading area into text information.

[0068] More specifically, step S31 includes:

[0069] S311, using a gyroscope sensor to obtain current screen posture information of the user terminal where the reading interface is located;

[0070] S312, determining the current sight line direction of the user using eye tracking technology, and locating the user's gaze position according to the current sight line direction;

[0071] S313: Determine a first reading area of the user according to the current screen posture information and the gaze position.

[0072] For example, in step S31, after the user enters the reading page of the electronic reading server, the eye tracking technology is used in combination with the gyroscope sensor to monitor the phone's posture information in real time to obtain the user's reading position and focus on the screen and capture the content the user is reading. When the user looks at a certain text content, the system automatically records and marks the part of the content as an important material for generating a video. Figure 3 As shown, Figure 3 This is a flowchart of a user reading content acquisition process provided by an embodiment of the present invention, eye position detection: using image processing and feature detection algorithms, such as Haar cascade detector or deep learning model, to identify the eye position in the user's facial image. These algorithms can locate the position of the eyes by detecting facial feature points or eye contours, and determine their coordinates in the image; eye movement tracking: using the changes between consecutive frame images, through optical flow method or feature point tracking and other methods, to track the movement of the eye on the screen. The optical flow method can infer the movement direction and speed of the eye based on the movement of pixels in the consecutive images, thereby determining the user's gaze point; eye movement model establishment: according to the eye movement law and the real-time data of eye tracking, an eye movement model is established. The model can include parameters such as gaze point, scanning path and dwell time. The model is updated and optimized through statistical analysis of the user's eye movement data; mobile phone gyroscope posture monitoring: the mobile phone's gyroscope sensor is used to monitor the user's mobile phone posture information. The gyroscope can detect the rotation direction and angle changes of the mobile phone, including the tilt and rotation of the mobile phone; gaze position and posture information fusion: the gaze position obtained by eye tracking and the posture information obtained by the mobile phone gyroscope monitoring are integrated, and the user's reading position on the screen is determined through comprehensive analysis based on the user's eye gaze position and the mobile phone's posture information; gaze position and text content association: the gaze position obtained by eye tracking is associated with the text content extracted from the screen, and a mapping relationship between the gaze position and the text is established. By comparing the coordinates of the gaze position and the text position, the text content that the user is gazing at is determined, thereby realizing the association between the gaze position and the text. In one specific embodiment, the above-described process can be expressed by calling the following formula: R = F(E, P, G). In this formula, R represents the content the user is currently reading; E represents the gaze position and eye movement data obtained by eye tracking; P represents the text content on the screen; and G represents the user's posture information obtained by the gyroscope. Function F represents the process of inferring the user's current reading content based on eye tracking data, on-screen text content, and user reading behavior. This function may include various data processing and analysis methods, such as text region identification, keyword matching, and matching gaze points with text regions. The specific function F can be designed and adjusted based on actual circumstances and may involve machine learning algorithms, pattern recognition techniques, natural language processing, and other methods.

[0073] More specifically, step S32 includes:

[0074] S321, obtaining current gesture touch data of the user on the user terminal;

[0075] S322, analyzing the current gesture touch data using gesture operation recognition technology to obtain the user's gesture operation and a movement trajectory of the gesture operation on the screen of the user terminal;

[0076] S323: Determine a second reading area of the user according to the gesture operation and the movement trajectory.

[0077] For example, in step S32, after the user enters the reading page of the electronic reading server, the system can analyze the user's gesture operations such as sliding and clicking with the help of gesture touch recognition technology to obtain the content that the user is paying attention to. When the user zooms in, shrinks, or annotates a certain text through gesture operations, the system will regard this part of the text content as the focus of the user's attention and use it as one of the important materials for generating a video. Figure 4 As shown, Figure 4This is another flowchart of a user reading content acquisition process provided by an embodiment of the present invention. Gesture touch data collection: When a user is reading an application or webpage, the system collects gesture touch data in real time through the touch screen of the mobile phone. The gesture touch data includes the user's gesture operations such as sliding, clicking, zooming in and out, as well as the position coordinates of the finger on the screen. Gesture recognition and analysis: Using machine learning or deep learning technology, the gesture touch data is recognized and analyzed. The system can identify different user gesture operations, such as single-finger sliding and two-finger zooming, as well as the movement trajectory of the finger on the screen. Key point extraction: Key points are extracted from the gesture touch data to determine the user's gaze position and focus area. Based on the trajectory and frequency of the gesture operation, the system extracts text areas and content that the user may be interested in. Text content extraction: Using optical character recognition (OCR) technology, the text content on the screen is extracted in real time, including the text area of the user's focus and the surrounding content. OCR technology converts the text on the screen into computer-recognizable text data for subsequent processing and analysis. Touch position and text content association: The key points extracted from the gesture touch data are associated with the extracted text content to establish a mapping relationship between the touch position and the text. The system identifies the text content currently being read by the user based on the user's gesture operation position. In a specific embodiment, the above-mentioned solution process can be expressed by calling the following formula: C=H(D,S), where: C represents the content that the user is currently reading; D represents gesture touch data, including the position of the finger on the screen, the type of gesture operation, etc.; S represents the text content on the screen; H represents a function that infers the content that the user is currently reading through gesture touch data and screen text content; function H can include a variety of data processing and analysis methods, such as gesture recognition, text area extraction, keyword matching, etc. The specific implementation method will depend on the specific situation and may involve machine learning algorithms, pattern recognition technology, natural language processing, etc.

[0078] Exemplarily, in step S33, the first reading area and the second reading area are merged into the current reading area of the user; or the overlapping area of the first reading area and the second reading area is used as the current reading area of the user.

[0079] Exemplarily, in step S34, the content of the current reading area is converted into text information using image processing and text recognition.

[0080] Specifically, step S4 includes:

[0081] S41, generating a video image based on the text information, the current emotion data, the current environmental factors, and the reading history data and inputting it into a trained multimodal fusion model;

[0082] S42, generating video content according to the text information;

[0083] S43, generating a presentation style of the video content according to the current emotion data, the current environmental factors and the reading history data;

[0084] S44, the video content and the presentation style generate a video picture.

[0085] Furthermore, the method further comprises:

[0086] The generated video images are edited and optimized to obtain the final video images.

[0087] Furthermore, the method further comprises:

[0088] Construct and train a multimodal fusion model to obtain a trained multimodal fusion model.

[0089] Exemplarily, first, data collection and preprocessing include user behavior data collection and data preprocessing; user behavior data collection: collect user behavior data on mobile reading applications, including read text content, eye tracking data, gesture touch data, emotion recognition data, etc. At the same time, collect user geographic location information and reading preference data, as well as environmental audio data and current time information; data preprocessing: clean and preprocess the collected data, remove noise data and outliers, perform preprocessing operations such as word segmentation and stop word removal on the text content, and standardize the geographic location information to make it suitable for model input.

[0090] Secondly, feature extraction and representation learning are performed: During the feature extraction and representation learning phase of text-to-video generation, data from multiple modalities is considered and converted into a unified feature representation to facilitate subsequent model training and generation. This process includes extracting text features, image features, emotion features, and ambient audio features. Text feature extraction: Convert the text content that the user is reading into a numerical vector. Here, the word embedding method is used to map each word into a low-dimensional continuous space so that the model can understand the semantic relationship between words. The word embedding models used include Word2Vec, GloVe and FastText. Image feature extraction: Extract high-level semantic features from the video frames obtained by eye tracking. Use pre-trained convolutional neural network models, such as ResNet, VGG or MobileNet, to extract the feature representation of the image. In the video frames obtained by eye tracking, the image is first pre-processed, scaled, cropped and normalized, and then forward propagated through the convolutional neural network model to obtain the feature vector of each video frame. Emotion feature extraction: Extract emotional information from the user's emotion recognition data. This is achieved through text sentiment analysis or facial expression recognition. For text sentiment analysis, models based on deep learning are used, such as Convolutional neural networks or recurrent neural networks are used to extract emotional information from text. For facial expression recognition, deep learning models or traditional image processing methods are used. Here, a face recognition model based on deep learning is used to identify the user's facial expressions and extract emotional features; environmental audio feature extraction: the process of extracting feature information related to the user's reading environment from environmental audio data. The feature extraction methods used here include spectrograms and Mel-frequency cepstral coefficients (MFCC). The process of converting audio signals into spectrograms in spectrograms can convert audio signals into spectral representations through Fourier transform, and then use filter banks to convert the spectrograms into logarithmic scale representations. MFCC is a feature extraction method for audio signal processing, which simulates the characteristics of the human auditory system. Therefore, we decompose the audio signal into a series of frequency band components and extract the Mel-frequency cepstral coefficients of each frequency band component as feature representations.

[0091] Next, we design and train the model: During the model design and training phase for text-generated videos, we need to design a multimodal fusion deep learning model that comprehensively leverages text features, image features, emotional features, and ambient audio features to generate personalized text-generated videos. This includes designing and training a multimodal fusion model. Multimodal fusion model design: Design an end-to-end deep learning model for comprehensive processing of multimodal information such as text, images, emotions, and audio. The overall architecture of the model can be divided into four main parts: text encoder, image encoder, emotion recognition module, and ambient audio processing module. The text encoder is responsible for encoding the text content that the user is reading into a fixed-length vector representation, using models such as recurrent neural networks (RNN) or Transformer; the image encoder is responsible for encoding the video frames obtained from eye tracking into a fixed-length vector representation, using a pre-trained convolutional neural network (CNN) model; the emotion recognition module is responsible for extracting emotional features from the user's emotion recognition data, using deep learning models such as CNN or RNN; the ambient audio processing module is responsible for extracting feature information related to the user's reading environment from the ambient audio data, using feature representations such as spectrograms or MFCC; Training strategy: Use an end-to-end approach to train the multimodal fusion model, adopt appropriate loss functions and optimization algorithms, and adopt multi-task learning or joint training methods, while considering the optimization objectives of multiple tasks, such as text generation, emotion recognition, and ambient audio classification.

[0092] Finally, model generation and editing are performed: After the model design and training are completed, the trained multimodal fusion model is used to generate personalized text-generated videos based on the user's input information, and the generated videos are edited to improve their viewing and appeal. Model generation and editing include personalized video generation and video editing and optimization. Personalized video generation: Using the trained multimodal fusion model, the user's current reading text content, past reading content, reading preferences, emotional state, ambient audio, and current time are input to generate personalized text-generated videos. The model can dynamically adjust the generated video content and style based on the user's reading habits and preferences, as well as the current mood and ambient audio, to enhance the user's reading experience; video editing and optimization: The generated text-generated videos may have some imperfections, which can be improved through video editing and optimization. For example, the video can be edited, the color and contrast can be adjusted, special effects and animations can be added, etc. The generated video can be personalized based on user feedback and preferences to meet user needs and expectations, thereby improving user satisfaction and user experience.

[0093] The following formula expresses the process of the above technical solution: Input data: text content currently read by the user: R, video frame sequence: C = (c1, c2, ..., cm), emotion recognition data: E, ambient audio data: A, previous reading content: P, reading preference: H;

[0094] Feature extraction and representation learning: text feature representation: Ftext(R), image feature representation: Fimage(C), emotion feature representation: Femotion(E), ambient audio feature representation: Faudio(A), past reading content feature representation: Fpast(P), reading preference feature representation: Fhabit(H);

[0095] Model design and training: Multimodal fusion model: Ffusion (Ftext, Fimage, Femotion, Faudio, Fpast, Fhabit), model parameters: m, loss function: L(Y, Y^);

[0096] Model generation and editing: generated text and video: Y = Fgeneration(Ffusion,m), edited video: Yedited = Fedit(Y).

[0097] In this process, feature extraction and representation learning are used to obtain feature representations for each input data point. These feature representations are then integrated using a multimodal fusion model to generate personalized text-generated videos. Finally, the generated videos are edited and optimized to enhance the user's reading experience.

[0098] Specifically, step S2 includes:

[0099] S21, obtaining the current facial information of the user;

[0100] S22, extracting facial expression features based on the current facial information;

[0101] S23, obtaining the user's current emotion data according to the facial expression features.

[0102] For example, Figure 5 As shown, Figure 5This is a block diagram of the user emotion recognition process provided by an embodiment of the present invention, including: facial detection: using image processing technology, such as a face detection algorithm based on deep learning, to detect the face area in the image captured by the camera; facial key point detection: for the detected face area, using a facial key point detection algorithm to locate the key feature points on the face, such as eyes, mouth, eyebrows, etc.; facial expression recognition: based on an existing facial expression database or using a deep learning model, the detected facial feature points are analyzed and identified to infer the user's emotional state. Common facial expressions include happiness, sadness, anger, surprise, etc.; emotional state classification: mapping the recognized facial expressions to predefined emotional state categories, such as "happy", "sad", "angry", etc. In a specific embodiment, the above-mentioned scheme process can be expressed by calling the following formula: face detection: F = FD(I), where I is the input image, FD is the face detection function, and F is the detected face area; key point detection: KP = KPD(F), where KPD is the key point detection function and KP is the detected facial key point; expression recognition: E = ER(KP), where ER is the expression recognition function and E is the inferred emotional state; emotion classification: CS = C(E), where C is the emotion classification function and CS is the result of mapping to the emotional state category.

[0103] In an embodiment of the present invention, the identified user emotional state can be used as one of the parameters for generating a video and passed to a video generation model. According to different emotional states, the style, theme, music and other content of the video can be adjusted to provide a more personalized video experience that conforms to the user's emotions.

[0104] For example, in step S1, Figure 6 As shown, Figure 6It is a block diagram of the environmental sound feature recognition process provided by an embodiment of the present invention, sound signal acquisition: turn on the mobile phone microphone and start recording environmental sound. The microphone will capture the sound signals in the surrounding environment, including background noise, human voice, environmental music, etc.; formula: A=MIC(I), where I is the instruction to start recording, MIC is the microphone acquisition function, and A is the recorded sound signal; sound signal preprocessing: preprocess the collected sound signal, including noise reduction processing, filtering processing, etc., to extract effective environmental sound information and reduce interference; formula: AP=PP(A), where PP is the preprocessing function of the sound signal, and AP is the preprocessed sound signal; environmental sound recognition: use machine learning or deep learning technology to analyze and recognize the preprocessed sound signal, and classify different types of environmental sounds, such as traffic noise, natural environmental sounds, human voices, etc.; formula: EC=ERC(AP), where ERC is the environmental sound recognition function, and EC is the recognized environmental sound category; near-field and far-field sound distinction: through sound The spectral characteristics and sound intensity of the signal are used to distinguish near-field sound from far-field sound; near-field sound refers to sound that is closer to the microphone, such as the user's own voice; and far-field sound refers to sound that is farther away from the microphone, such as other sounds in the environment; Formula: NF = NFS (AP), where NFS is the near-field sound recognition function and NF is the recognized near-field sound; Ambient sound parameter extraction: Based on the recognized ambient sound type and near-field / far-field sound information, the characteristic parameters of the ambient sound are extracted, such as audio spectrum, sound intensity, sound rhythm, etc.; Formula: EP = EPP (EC, NF), where EPP is the ambient sound parameter extraction function and EP is the extracted ambient sound parameter; Ambient sound parameter output: The extracted ambient sound parameter is used as one of the parameters for generating the video and passed to the video generation model. According to different ambient sounds, the style, theme, sound effects and other content of the video are adjusted to provide a more personalized video experience that conforms to the current environment; Formula: P = EP, the ambient sound parameter is used as the input parameter of the video generation model. The embodiment of the present invention can use the mobile phone microphone to obtain current ambient sound information and use it as one of the important parameters for generating video, thereby providing users with a richer and more personalized video experience.

[0105] like Figure 7 As shown, Figure 7This is another block diagram of a video image generation process provided by an embodiment of the present invention. After a user enters a reading page on an electronic reading server, the user's eye information is captured by a camera, eyeball characteristics are analyzed, and the user's line of sight direction is determined. The phone's posture is monitored using a gyroscope sensor to determine the phone's tilt angle and direction. The line of sight and posture are combined to determine the screen area LA that the user is focusing on. The finger operation and movement trajectory are analyzed to determine the screen area LB that the user is focusing on. Combining LA and LB, the recognition range LC is determined. A screenshot or screen recording is performed, and the current screen content is obtained through image processing and text recognition. Content is extracted based on the recognition range LC to obtain content CA. The current time and geographic location are obtained in real time. The user's facial information is captured by a camera, facial expression features are extracted, and the user's current emotional attributes are determined. The ambient sound information is captured by a microphone, ambient sound features are extracted, and the attributes of the ambient sound are determined. Past browsing history and interactive behavior data (likes, favorites, comments, etc.) are collected. The content CA, current time, geographic location, user's current emotional attributes, ambient sound attributes, past browsing history, and interactive behavior data are input into a multimodal fusion model to generate a personalized video, which can be edited, created, and shared.

[0106] The embodiment of the present invention discloses a method for generating a video image. The method obtains a user's current environmental factors and reading history data; identifies the user's current emotional data; locates the user's current reading area and obtains text information in the current reading area; and generates a video image based on the text information, the current emotional data, the current environmental factors, and the reading history data. The method can generate a personalized video image based on the user's current reading text information, current emotional data, current environmental factors, and reading history data. The method uses video visuals to assist the user in reading, resolving issues such as limited imagination, subjective understanding differences, and an obscure reading experience during the user's reading process. The method makes reading more vivid and interesting, enhancing the user's reading experience.

[0107] See also Figure 8 , Figure 8 1 is a structural diagram of a video image generating device 10 provided by an embodiment of the present invention. The video image generating device 10 includes:

[0108] Related data acquisition module 11, used to obtain the user's current environment factors and reading history data;

[0109] An emotion data recognition module 12 is used to recognize the current emotion data of the user;

[0110] A text information acquisition module 13 is used to locate the current reading area of the user and obtain text information of the current reading area;

[0111] The video image generation module 14 is used to generate a video image according to the text information, the current emotion data, the current environmental factors and the reading history data.

[0112] Furthermore, the video image generating device 10 further includes:

[0113] The video image presentation module is used to synchronously present the video image above (or below) the current reading area.

[0114] Specifically, the text information acquisition module 13 is used to:

[0115] Determining a first reading area of the user according to current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user;

[0116] determining a second reading area of the user according to current gesture touch data of the user on the user terminal;

[0117] determining a current reading area of the user according to the first reading area and the second reading area;

[0118] The content of the current reading area is converted into text information.

[0119] More specifically, determining the first reading area of the user according to the current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user includes:

[0120] A gyroscope sensor is used to obtain the current screen posture information of the user terminal where the reading interface is located;

[0121] Determining the user's current line of sight using eye tracking technology, and locating the user's gaze position based on the current line of sight;

[0122] A first reading area of the user is determined according to the current screen posture information and the gaze position.

[0123] More specifically, determining the second reading area of the user according to the current touch gesture data of the user on the user terminal includes:

[0124] Acquiring current gesture touch data of the user on the user terminal;

[0125] Analyzing the current gesture touch data using gesture operation recognition technology to obtain the user's gesture operation and a movement trajectory of the gesture operation on the screen of the user terminal;

[0126] A second reading area of the user is determined according to the gesture operation and the movement trajectory.

[0127] Specifically, the video image generation module 14 is used to:

[0128] The text information, the current emotion data, the current environmental factors and the reading history data are used to generate a video image and input it into the trained multimodal fusion model;

[0129] generating video content according to the text information;

[0130] generating a presentation style of the video content according to the current emotion data, the current environmental factors, and the reading history data;

[0131] The video content and the presentation style generate a video picture.

[0132] Furthermore, the device further comprises:

[0133] The video editing module is used to edit and optimize the generated video images to obtain the final video images.

[0134] Furthermore, the device further comprises:

[0135] The model building module is used to build and train the multimodal fusion model to obtain a trained multimodal fusion model.

[0136] Specifically, the emotion data recognition module 12 is used to:

[0137] Obtaining the current facial information of the user;

[0138] extracting facial expression features based on the current facial information;

[0139] The user's current emotion data is obtained according to the facial expression features.

[0140] A video image generating device 10 provided in an embodiment of the present invention can implement all processes of the video image generating method of the above embodiment. The functions of each module in the device and the technical effects achieved are respectively the same as the functions and technical effects achieved by the video image generating method of the above embodiment, and will not be repeated here.

[0141] See also Figure 9 , is a schematic diagram of the structure of a video image generation device 20 provided in an embodiment of the present invention. The video image generation device 20 of this embodiment includes a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above-described video image generation method embodiment are implemented. Alternatively, when the processor 21 executes the computer program, the functions of the modules in the above-described video image generation device embodiment are implemented.

[0142] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory 22 and executed by the processor 21 to implement the present invention. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the video image generation device 20.

[0143] The video image generation device 20 can be a computing device such as a desktop computer, laptop, PDA, or cloud server. The video image generation device 20 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will appreciate that the schematic diagram is merely an example of the video image generation device 20 and does not limit the video image generation device 20. The video image generation device 20 may include more or fewer components than shown, or may combine certain components or different components. For example, the video image generation device 20 may also include input and output devices, network access devices, buses, and the like.

[0144] The processor 21 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. The processor 21 is the control center of the video image generation device 20 and connects various parts of the entire video image generation device 20 using various interfaces and lines.

[0145] The memory 22 can be used to store the computer programs and / or modules. The processor 21 implements the various functions of the video image generation device 20 by running or executing the computer programs and / or modules stored in the memory 22 and accessing the data stored in the memory 22. The memory 22 can mainly include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 22 can include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0146] Wherein, if the module integrated in the video image generating device 20 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0147] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0148] An embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the video image generation method as described in the above embodiment.

[0149] An embodiment of the present invention further provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the video image generation method of the above embodiment.

[0150] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for generating a video image, characterized in that: include: Obtain the user's current environment and reading history data; identifying current emotion data of the user; Locating the current reading area of the user and obtaining text information of the current reading area; A video image is generated according to the text information, the current emotion data, the current environmental factors and the reading history data.

2. The video image generation method according to claim 1, wherein: The locating the current reading area of the user and obtaining text information of the current reading area includes: Determining a first reading area of the user according to current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user; determining a second reading area of the user according to current gesture touch data of the user on the user terminal; determining a current reading area of the user according to the first reading area and the second reading area; The content of the current reading area is converted into text information.

3. The video image generation method according to claim 2, wherein: The determining of the first reading area of the user according to the current screen posture information of the user terminal where the reading interface is located and the current sight direction of the user includes: A gyroscope sensor is used to obtain the current screen posture information of the user terminal where the reading interface is located; Determining the user's current line of sight using eye tracking technology, and locating the user's gaze position based on the current line of sight; A first reading area of the user is determined according to the current screen posture information and the gaze position.

4. The video image generation method according to claim 2, wherein: The determining the second reading area of the user according to the current gesture touch data of the user on the user terminal includes: Acquiring current gesture touch data of the user on the user terminal; Analyzing the current gesture touch data using gesture operation recognition technology to obtain the user's gesture operation and a movement trajectory of the gesture operation on the screen of the user terminal; A second reading area of the user is determined according to the gesture operation and the movement trajectory.

5. The video image generation method according to claim 1, wherein: Generating a video image according to the text information, the current emotion data, the current environmental factors and the reading history data includes: The text information, the current emotion data, the current environmental factors and the reading history data are used to generate a video image and input it into the trained multimodal fusion model; generating video content according to the text information; generating a presentation style of the video content according to the current emotion data, the current environmental factors, and the reading history data; The video content and the presentation style generate a video picture.

6. The video image generation method according to claim 1, wherein: The identifying the current emotion data of the user includes: Obtaining the current facial information of the user; extracting facial expression features based on the current facial information; The user's current emotion data is obtained according to the facial expression features.

7. A video image generating device, characterized in that: include: Related data acquisition module, used to obtain the user's current environment factors and reading history data; An emotion data recognition module, configured to recognize the user's current emotion data; A text information acquisition module, configured to locate the current reading area of the user and obtain text information of the current reading area; The video frame generation module is used to generate a video frame according to the text information, the current emotional data, the current environmental factors and the reading history data.

8. A video image generating device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for generating a video picture according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the video image generation method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the video picture generation method according to any one of claims 1 to 6.