Data Processing Method, Apparatus, Electronic Device, and Storage Medium

The data processing method and apparatus address the limitations of conventional teleprompters by processing audio-video frame data to adjust the target user's gaze and synchronize audio information with text, resulting in a more versatile and efficient broadcasting solution.

JP7691055B2Active Publication Date: 2025-06-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023578190
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-31
Filing Date
2022-08-24
Publication Date
2025-06-11
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

Conventional teleprompters are cumbersome, occupy a large area, and lack versatility, making them inconvenient for modern video creators who need flexible and efficient broadcasting solutions.

Method used

A data processing method and apparatus that collect audio-video frame data from a target user, process the face image to adjust the gaze angle, and synchronize the audio information with the target text, allowing for real-time display of the target statement and face image on a client device.

Benefits of technology

This solution enhances broadcasting convenience and versatility by reducing the physical space required and simplifying operations, enabling efficient broadcasting with a mobile terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691055000001
    Figure 0007691055000001
  • Figure 0007691055000002
    Figure 0007691055000002
  • Figure 0007691055000003
    Figure 0007691055000003
Patent Text Reader

Abstract

This specification discloses a data processing method, an apparatus, an electronic device and a storage medium, which includes: collecting audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed and a facial image to be processed, processing the facial image to be processed according to a target gaze angle adjustment model to obtain a target facial image corresponding to the facial image to be processed, tracking and processing the audio information to be processed according to an audio content tracking method to determine a related target statement in a target text of the audio information to be processed, respectively displaying the target statement and the target facial image to a client related to the target user, or simultaneously displaying the target statement and the target facial image to a client related to the target user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Related Application] This application was filed with the China National Intellectual Property Administration on August 31, 2021, claiming the priority of a Chinese patent application with the application number 202111016229.0, the entire content of which is incorporated herein by reference.

[0002] This disclosure relates to the field of computer technologies, for example, data processing methods, apparatuses, electronic devices, and storage media.

Background Art

[0003] Teleprompters can be used in various broadcast scenarios. A teleprompter may display speech text for a speaker when the speaker is making a speech. Usually, the teleprompter is arranged below the front of the camera lens, and between the teleprompter and the camera lens, a transparent glass piece or a special beam splitter is arranged at an angle of 45 degrees. This glass or beam splitter is provided to reflect light from the direction of the teleprompter to the speaker and transmit light from the direction of the speaker to the camera lens. Further, a light shield provided around the camera lens in conjunction with the back side of the glass or beam splitter also prevents unnecessary reflection of light onto the camera lens.

[0004] With the development of Internet technologies, each video creator may become a user of a teleprompter. However, conventional teleprompters have problems such as a large occupied area, complex operation procedures, and poor versatility.

Summary of the Invention

[0005] This disclosure provides a data processing method, apparatus, electronic device, and storage medium for achieving the technical effects of broadcast convenience and versatility.

[0006] This disclosure provides a data processing method, the method comprising Collecting audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed and a face image to be processed, Processing the face image to be processed based on a target gaze angle adjustment model to obtain a target face image corresponding to the face image to be processed, Performing follow-up processing on the audio information to be processed based on an audio content following method to determine a related target statement in the target text of the audio information to be processed, Displaying the target statement and the target face image on a client related to the target user respectively, or simultaneously displaying the target statement and the target face image on a client related to the target user, including.

[0007] The present disclosure further provides a data processing device, and the device includes, An audio-video frame data collection module configured to collect audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed and a face image to be processed, A face image processing module configured to process the face image to be processed based on a target gaze angle adjustment model to obtain a target face image corresponding to the face image to be processed, A target statement determination module configured to perform follow-up processing on the audio information to be processed based on an audio content following method to determine a related target statement in the target text of the audio information to be processed, A display module configured to display the target statement and the target face image on a client related to the target user respectively, or simultaneously display the target statement and the target face image on a client related to the target user, and includes.

[0008] The present disclosure further provides an electronic device, and the electronic device includes, One or more processors, A storage device storing one or more programs, and when the one or more programs are executed by the one or more processors, causing the one or more processors to execute the data processing method described above.

[0009] The present disclosure further provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, cause the data processing method described above to be executed.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, the present disclosure can be implemented in various forms, and these embodiments are used for the purpose of understanding the present disclosure. The accompanying drawings and embodiments of the present disclosure are for illustrative purposes only.

[0012] It should be understood that the multiple steps described in the method embodiments of the present disclosure may be executed in a different order and / or in parallel. Furthermore, method embodiments can add steps and / or omit the steps shown. The scope of the present disclosure is not particularly limited in this regard.

[0013] The term "comprising" and its variants used herein are open-ended, that is, they mean "including but not limited to". The term "based on" means "based at least in part on". The term "an embodiment" means "at least one embodiment", the term "another embodiment" means "at least one other embodiment", and the term "some embodiments" means "at least some embodiments". Related definitions of other terms are shown in the following description.

[0014] Concepts such as "first", "second", etc. referred to in the present disclosure are used to distinguish different devices, modules or units, and are not used to define the order or interdependence of the functions executed by these devices, modules or units. It should be noted that the modifications of "one" and "a plurality" referred to in the present disclosure are illustrative rather than limiting, and those skilled in the art should understand that "one or a plurality" should be understood unless explicitly stated in the context.

[0015] The names of messages or information that interact between multiple devices in the embodiments of the present disclosure are used only for illustrative purposes and are not intended to limit the scope of these messages or information.

[0016] Example 1 FIG. 1 is a schematic flowchart of a data processing method provided by Embodiment 1 of the present disclosure. Embodiments of the present disclosure are applicable to the scenario of audio broadcasting of a target user, and while dynamically following the target text for the user, process the user's face image. The method can be executed by a data processing device, which can be implemented by software and / or hardware, for example, may be implemented by an electronic device, and the electronic device may be a mobile terminal, a personal computer (PC) terminal, a server, or the like. The real-time interaction application scenario can usually be realized by the cooperation of a client and a server. The method provided by this embodiment may be executed by the client, may be executed by the server, or may be executed by both in cooperation.

[0017] S110. Collect audio-video frame data related to the target user.

[0018] The technical solution provided by the embodiment of the present disclosure may be integrated into one application, develop an application based on the technical solution of the present disclosure, and the user may perform audio broadcasting, video broadcasting, etc. based on the application.

[0019] Audio-video frames may be collected when a user interacts through a real-time interaction interface, and the real-time interaction interface may be any interaction interface in a real-time interaction application scenario. The real-time interaction scenario may be realized by the Internet and computer means, for example, it may be an interaction application realized by a native program or a web program, etc. The real-time interaction application scenario may be a live broadcast scenario, a video conference scenario, an audio broadcast scenario, and a scenario of a recorded video. The live broadcast scenario may include a live broadcast of product sales in an application and a live broadcast scenario based on a live broadcast platform. The audio broadcast scenario may be a scenario in which a caster of a TV station broadcasts corresponding content and packages the multimedia data stream broadcast by the caster by a camera to at least one client. The audio-video frame data includes audio information to be processed and face images to be processed. When a target user broadcasts based on target text, the audio collection device may collect the audio information of the target user and use the collected audio information as the audio information to be processed. For example, the audio collection device may be a microphone array on a mobile terminal or a microphone array provided in the environment where the broadcasting user is located. Correspondingly, during the user's broadcasting process, the shooting device may collect the face image information of the target user, and at this time, the collected image may also be used as the face image to be processed.

[0020] In this embodiment, as a method for collecting audio-video frame data related to the target user, it may be collected in real time or periodically. For example, a real-time interaction interface is generated based on an online video broadcast scenario. The video broadcast includes a caster (target user) and viewers who watch the caster's broadcast. When the caster broadcasts based on a preset broadcast statement, the imaging device and the audio collection device collect the audio information and face image information of the target user in real time or every few seconds, for example, every 5 seconds, to obtain audio-video frame data.

[0021] When the user interacts based on the real-time interaction interface, the imaging device and the audio collection device collect the face image information and audio information corresponding to the target user, process the audio information and the face image, and transmit the processed face image and audio information as a data stream to other clients.

[0022] In this embodiment, when the target user broadcasts content, the target user mainly broadcasts based on the target text, and the target text may be uploaded to an application developed based on the present technical solution, or uploaded to the prompt program of the mobile terminal, and the target text may be processed by the application or the prompt program.

[0023] In this embodiment, collecting the audio-video frame data related to the target user includes, when it is detected that a preset event is triggered, collecting the audio information to be processed of the target user based on the audio collection device, and collecting the face image to be processed of the target user based on the imaging device.

[0024] The preset event may be the triggering of a wake-up word, the triggering of a gaze adjustment control, the detection that the user is present in front of the display screen, and / or the detection that audio information has been collected.

[0025] In the application process, when it is detected that the target user has triggered a preset event, the audio information to be processed and the face image to be processed of the target user are collected.

[0026] S120. Process the face image to be processed based on the target line-of-sight angle adjustment model to obtain a target face image corresponding to the face image to be processed.

[0027] The photographing device on a general terminal device is mostly installed at a specific position of the terminal. For example, the photographing device is installed at the top edge position of the terminal. When the target user interacts with other users based on the display screen, if there is a certain deviation between the photographing angle of the photographing device and the line of sight of the target user, the line of sight of the target user may not be on the same horizontal line as the photographing device, and there may be a certain angular deviation in the line-of-sight angle of the target user that the user watching the live broadcast sees. For example, there is a situation where the viewing user looks obliquely at the line of sight of the target user with the terminal. Or, when the target user broadcasts with reference to the target text displayed on the teleprompter, when the target user needs to look at the target text, due to the positional deviation between the target text display position and the photographing device, a situation where the line of sight of the target user is not in focus occurs, which has an adverse effect on the interaction effect. To solve the above problems, the line-of-sight angle of the user in the face image to be processed can be adjusted by the target line-of-sight angle adjustment model.

[0028] The target gaze angle adjustment model is a pre-trained model and is used to focus on the gaze of the target user in the face image to be processed. The target face image is an image obtained by performing a focusing process on the gaze in the face image to be processed by the target gaze angle adjustment model. Through the focusing process, the user's gaze angle in the face image to be processed is focused on the target angle. For example, the pupil position of the eyeball is adjusted so that the gaze angle of the target user can be adjusted to the target gaze angle. The target gaze angle, also referred to as the target angle, is the angle at which the user's gaze is perpendicular to the display screen, that is, the angle at which the user's gaze directly looks at the display screen. The target gaze angle may be any pre-set angle. In order to improve the interaction efficiency between the caster user and other users, the target gaze angle may be the angle at which the gaze of the target user and the imaging device are on the same horizontal line.

[0029] After collecting the face image to be processed, the face image to be processed is input into the target gaze angle adjustment model to perform an adjustment process, and the gaze angle of the target user in the face image to be processed is adjusted to the target angle.

[0030] The gaze of the target user in the face image to be processed may or may not be focused. In fact, in order to avoid processing all the face images to be processed, after obtaining the face image to be processed, the face image to be processed may be pre-processed.

[0031] In one embodiment, based on the feature detection module, it is determined whether the gaze feature in the face image to be processed matches the preset gaze feature. If the gaze feature in the face image to be processed does not match the preset gaze feature, the face image to be processed is processed based on the target gaze angle adjustment model to obtain the target face image.

[0032] The feature detection module detects the user's gaze feature and is mainly used to determine whether the user's gaze angle matches the target angle phase. The preset gaze feature is a feature that matches the target angle. The preset gaze feature may be a feature such as eyelids and pupils, for example, the position of the pupil in the eye socket.

[0033] After acquiring the face image to be processed, the face image to be processed is processed based on the feature detection module, and it is determined whether the gaze feature in the face image to be processed matches the preset gaze feature. If the gaze feature in the face image to be processed does not match the preset gaze feature, the gaze angle of the target user does not match the target angle. At this time, the face image to be processed is subjected to focusing processing based on the target gaze angle adjustment model to obtain the target face image.

[0034] S130. Follow-up process the audio information to be processed based on the audio content follow-up method, and determine the relevant target statement in the target text of the audio information to be processed.

[0035] The audio content follow-up method is a method of identifying the audio information of the target user during the broadcast process of the target user, determining the position of the broadcast content in the target text, and following the text content of the voice broadcast in real time. That is, regardless of whether the voice broadcast speed of the target user is fast or slow, it can follow according to the real-time speech speed of the target user, and solve the problem that the speed of the target user's broadcast text becomes slower or faster than the text scroll speed due to a fixed scroll speed. The target statement is the statement in the target text corresponding to the audio information to be processed. The target text is the text that needs to be uploaded in advance by the user and broadcast according to its content.

[0036] The audio information to be processed can be identified and followed up according to the preset audio content follow-up method, and the corresponding target statement in the target text of the audio information to be processed can be determined.

[0037] S140. Display the target statement and the target face image on the client related to the target user respectively, or display the target statement and the target face image on the client related to the target user at the same time.

[0038] Displaying the target statement and the target face image to the associated clients respectively may also mean displaying the target statement to some clients and the target face image to other clients. Correspondingly, displaying the target statement and the target face image to the associated clients simultaneously may mean that, for the clients associated with the target user, both the target statement and the target face image may be displayed to such clients. Those skilled in the art should understand that the differentiated display of the target statement and the target face image can be set according to actual needs and is not particularly limited in the embodiments of the present disclosure.

[0039] During the display process, the target statement may be displayed in a distinguishable manner. Here, distinguishable display means differentiating the display method of the target statement from the display methods of other contents of the target text. For example, the target statement may be associated with the content spoken by the target user and displayed alone on the client, or the target statement and the statements before and after the target statement may be displayed on the client simultaneously, and the target statement may be highlighted by bold font display, emphasis display, etc. The multimedia data stream may be a data stream generated during the broadcast of the target user, for example, all video streams during the broadcast. When at least one user watches the target user broadcast via a terminal device, the client corresponding to the viewing user may be a client associated with the target user. The target face image and the audio information may be transmitted as a multimedia data stream to at least one client, and the target face image of the target user and the audio information corresponding to the broadcast target user may be displayed on the client.

[0040] Finally, what is displayed on the client is the video stream corresponding to the target user broadcast.

[0041] After determining the target statement, the target statement is separately displayed as the target text, and at this time, the target text is displayed on the terminal device corresponding to the target user. At the same time, the target face image corresponding to the target user is sent to the client related to the target user, and the target user displayed on the client corresponding to the viewing user becomes the image after focusing by the target line-of-sight angle adjustment model.

[0042] To understand the technical effect of the technical solution of the present disclosure, referring to the schematic diagram shown in FIG. 2, if the line of sight of the user in the detected face image to be processed is inclined, after the focusing process by the target line-of-sight angle adjustment model, the front view image shown in FIG. 2 is obtained, and the line-of-sight angle in the front view image coincides with the target angle. The image displayed on the client is the target face image after line-of-sight focusing.

[0043] According to the technical solution of the embodiment of the present disclosure, the audio information to be processed and the face image to be processed of the target user are collected, and further, the audio information to be processed and the face image to be processed are processed to obtain the target face image and the position of the target text of the target statement corresponding to the audio information to be processed, solving the problems that the occupied area of the teleprompter in the related art is large, the use is cumbersome, and the use of the teleprompter is inconvenient, and realizing the technical effect of realizing efficient broadcasting only with a mobile terminal.

[0044] Embodiment 2 FIG. 3 is a schematic flowchart of the data processing method provided by Embodiment 2 of the present disclosure. Based on the above embodiment, before processing the face image to be processed based on the target line-of-sight angle adjustment model, the target line-of-sight angle adjustment model can be obtained by training. For the specific implementation form of obtaining the target line-of-sight angle adjustment model by training, reference may be made to the description of the technical solution of the present disclosure. Here, the description of the same or corresponding technical terms as the above embodiment will not be repeated.

[0045] As shown in FIG. 3, the method includes the following steps.

[0046] S210. Obtain a training sample set.

[0047] Before obtaining a target line-of-sight angle adjustment model through training, it is first necessary to obtain training samples and train based on the training samples. In order to improve the accuracy of the model, as many and as rich training samples as possible can be obtained.

[0048] The training sample set includes a plurality of training samples. Each training sample includes a target line-of-sight angle image and a non-target line-of-sight angle image. The training sample is determined based on a target sample generation model obtained through pre-training.

[0049] The target line-of-sight angle image is a face image in which the user's line of sight coincides with the target angle. The non-target line-of-sight angle image is a face image in which the user's line of sight does not coincide with the target angle. The target sample generation model may be a model that generates training samples.

[0050] First, obtain a target sample generation model through training. The target sample generation model includes a positive sample generation sub-model and a negative sample generation sub-model. The positive sample generation sub-model is used to generate a target line-of-sight angle image in the training sample image. The user's line of sight angle in the target line-of-sight angle image coincides with the target angle. Correspondingly, the negative sample generation sub-model is used to generate a non-target line-of-sight angle image in the training sample image, and the user's line of sight angle in the non-target line-of-sight angle image does not coincide with the target angle.

[0051] S220. For each training sample, input the non-target line-of-sight angle image in the current training sample into the line-of-sight angle adjustment model to be trained, and obtain the actual output image corresponding to the current training sample.

[0052] S230. Based on the actual output image and the target line-of-sight angle image of the current training sample, determine a loss value, and based on the loss value and the preset loss function of the line-of-sight angle adjustment model to be trained, adjust the model parameters of the line-of-sight angle adjustment model to be trained.

[0053] S240. Converge with the preset loss function as the training target to obtain the target line-of-sight angle adjustment model.

[0054] Regarding the determination process of the above target line-of-sight angle adjustment model, train the line-of-sight angle adjustment model to be trained based on each training sample in the training sample set to obtain the target line-of-sight angle adjustment model. By using each non-target line-of-sight angle image in the training sample as the input of the line-of-sight angle adjustment model to be trained and the target line-of-sight angle image corresponding to the non-target line-of-sight angle image as the output of the line-of-sight angle adjustment model to be trained, adjust the model parameters in the line-of-sight angle adjustment model to be trained. When the convergence of the loss function in the line-of-sight angle adjustment model to be trained is detected, it is determined that the target line-of-sight angle adjustment model is obtained through training.

[0055] S250. Collect audio-video frame data related to the target user, where the audio-video frame data includes the audio information to be processed and the face image to be processed.

[0056] S260. Process the face image to be processed based on the target line-of-sight angle adjustment model to obtain the target face image corresponding to the face image to be processed.

[0057] S270. Follow and process the audio information to be processed based on the audio content following method to determine the target statement in the target text of the audio information to be processed.

[0058] S280. Display the target statement and the target face image on the client related to the target user respectively, or display the target statement and the target face image on the client related to the target user simultaneously.

[0059] In this embodiment, displaying the target statement and the target face image to the client associated with the target user respectively or simultaneously means that the target statement is separately displayed as the target text on the first client, and the target audio-video frame corresponding to the target face image is displayed on the second client.

[0060] The first client and the second client are relative and may be the same or different. The first client may be the client used by the target user, and the second client may be the client used by other users who watch the live broadcast of the target user.

[0061] In the case of a real-time interaction scenario, each face image to be processed that is collected may be processed, and the obtained target face image may be transmitted to other clients in the form of a multimedia data stream. Thereby, not only can the captured video be made more dynamic and interactive, but also the viewing user can view the image with the line of sight always focused on the target angle, achieving the effect of improving the user's viewing experience.

[0062] According to the technical solution of the embodiment of the present disclosure, before processing the face image to be processed based on the target line-of-sight angle adjustment model, first obtain the target line-of-sight angle adjustment model through training, process each face image to be processed collected by the imaging device based on the target line-of-sight angle adjustment model to obtain a target face image with the line of sight focused, transmit the target face image to at least one client, and each user can view the target user after the line of sight is focused, and a more interactive video stream can be obtained.

[0063] Example 3 Figure 4 is a schematic flowchart of the data processing method provided by Embodiment 3 of the present disclosure. Based on the foregoing embodiments, before obtaining a target line-of-sight angle adjustment model through training, corresponding training samples are generated based on a target sample generation model. Correspondingly, before obtaining the training samples, first train to obtain a target sample generation model. For specific embodiments, reference may be made to the description of the technical solution of the present disclosure. Here, descriptions of the same or corresponding technical terms as those in the above embodiments will not be repeated.

[0064] As shown in Figure 4, the method includes the following steps.

[0065] S310. Train to obtain a non-target line-of-sight angle image generation sub-model in the target sample generation model.

[0066] Input a pre-collected Gaussian distribution vector and an original non-frontal view sample image into the non-target line-of-sight angle image generation sub-model to be trained to obtain an error. Based on the error and the loss function in the non-target line-of-sight angle image generation sub-model to be trained, correct the model parameters in the non-target line-of-sight angle image generation sub-model to be trained, converge the loss function as the training target, obtain the non-target line-of-sight angle image generation sub-model, and generate non-target line-of-sight angle images in the training samples based on the non-target line-of-sight angle image generation sub-model.

[0067] In this embodiment, inputting the pre-collected Gaussian distribution vector and the original non-frontal view sample image into the non-target line-of-sight angle image generation sub-model to be trained to obtain an error means processing the Gaussian distribution vector based on the generator in the non-target line-of-sight angle image generation sub-model to be trained to obtain an image to be compared, and processing the original non-frontal view sample image and the image to be compared based on the discriminator in the non-target line-of-sight angle image generation sub-model to be trained to obtain the error.

[0068] The Gaussian distribution vector may be random sampling noise. When the user is not looking directly ahead, the face image is collected to obtain the original non-frontal view sample image. The model parameters in the non-target line-of-sight angle image generation sub-model to be trained are default parameter values. The actual output result is obtained by using the Gaussian distribution vector and the original non-frontal view sample image as the input of the non-target line-of-sight angle image generation sub-model to be trained. An error can be obtained based on the actual output result and the original non-frontal view sample image. Based on the error and the preset loss function in the non-target line-of-sight angle image generation sub-model, the model parameters in the sub-model can be corrected. The non-target line-of-sight angle image generation sub-model can be obtained by converging the loss function as the training target.

[0069] When training multiple sub-models of the present technical solution disclosure, adversarial training may be used. As adversarial training, a generator and a discriminator are included in the non-target line-of-sight angle image generation sub-model. The generator is used to process the Gaussian distribution vector to generate a corresponding image. The discriminator determines the similarity between the generated image and the original image, and is used to adjust the model parameters in the generator and the discriminator based on the error until the training of the non-target line-of-sight angle image generation sub-model is completed.

[0070] The generator in the non-target line-of-sight angle image generation sub-model processes the Gaussian distribution vector to obtain an image to be compared corresponding to the Gaussian distribution vector. At the same time, the image to be compared and the original non-frontal view sample image are input into the discriminator, and the discriminator can perform a determination process on the two images to obtain an output result. The model parameters in the generator and the discriminator can be corrected according to the output result. When the convergence of the loss function of the model is detected, the obtained model can be used as the non-target line-of-sight angle image generation sub-model.

[0071] Obtain the model parameters in the non-target line-of-sight angle image generation sub-model, multiplex them into the target line-of-sight angle image generation sub-model to be trained, train the target line-of-sight angle image generation sub-model to be trained based on the pre-collected Gaussian distribution vectors and the original front-view sample images, and obtain the target line-of-sight angle image generation sub-model.

[0072] After obtaining the non-target line-of-sight angle generation sub-model, obtain the target line-of-sight angle generation sub-model through training. For example, obtain the model parameters in the non-target line-of-sight angle image generation sub-model, multiplex them into the target line-of-sight angle image generation sub-model to be trained, and train the target line-of-sight angle image generation sub-model to be trained based on the pre-collected Gaussian distribution vectors and the original front-view sample images to obtain the target line-of-sight angle image generation sub-model.

[0073] At this time, the target line-of-sight angle image generation sub-model to be trained is also completed based on adversarial training. That is, the sub-model also includes a generator and a discriminator. The functions of the generator and the discriminator are the same as those of the above sub-model, and the method for obtaining the target line-of-sight angle image generation sub-model through training is the same as the method for obtaining the non-target line-of-sight angle image generation sub-model, so the description here will not be repeated.

[0074] To improve the convenience of obtaining the target line-of-sight angle image generation sub-model through training, after the training of the non-target line-of-sight angle image generation sub-model is completed, multiplex the model parameters in the non-target line-of-sight angle image generation sub-model and use it as the initial model parameters in the target line-of-sight angle image generation sub-model obtained through training.

[0075] S330: Input each of the multiple Gaussian distribution vectors to be trained into the target line-of-sight angle image generation sub-model and the non-target line-of-sight angle image generation sub-model to obtain the target line-of-sight angle images and non-target line-of-sight angle images in the training samples.

[0076] The entire target line-of-sight angle image generation sub-model and non-target line-of-sight angle image generation sub-model may be used as a target sample generation model. The target line-of-sight angle image generation sub-model and non-target line-of-sight angle image generation sub-model may be encapsulated into capsules, and two images can be output based on the input. At this time, the line-of-sight angles of the user in the two images are different.

[0077] A common problem with training models is that a large number of samples need to be collected, and there is a certain degree of difficulty in collecting samples. For example, for the images of a large number of users under the target line-of-sight angle and non-target line-of-sight angle in this embodiment, it is difficult to collect samples, and there is a problem that the specifications do not match. Based on this technical solution, random sampling noise is directly processed to obtain images of the same user under different line-of-sight angles, obtain corresponding samples, improve the convenience and versatility of sample determination, and further improve the convenience of the training model.

[0078] In one embodiment, based on the target line-of-sight angle image generation sub-model and non-target line-of-sight angle image generation sub-model in the target sample generation model, a plurality of Gaussian distribution vectors are sequentially processed to obtain the target line-of-sight angle image and non-target line-of-sight angle image in the training samples.

[0079] S340. Obtain a target line-of-sight angle adjustment model through training based on a plurality of training samples.

[0080] According to the technical solution of the embodiment of the present disclosure, the target sample generation model obtained through pre-training can process random sampling noise to obtain a large number of training samples for training the target line-of-sight angle adjustment model, improving the convenience and uniformity of obtaining training samples.

[0081] Example 4 FIG. 5 is a schematic flowchart of the data processing method provided by Embodiment 4 of the present disclosure. Based on the foregoing embodiments, the audio information to be processed is tracked based on a preset audio content tracking method, and the target statement in the target text of the audio information to be processed is determined. For specific embodiments, reference may be made to the description of the technical solution of the present disclosure. Here, descriptions of the same or corresponding technical terms as those in the above embodiments will not be repeated.

[0082] As shown in FIG. 4, the method includes the following steps.

[0083] S410. Collect audio-video frame data related to the target user. Here, the audio-video frame data includes the audio information to be processed and the face image to be processed.

[0084] S420. Process the face image to be processed based on the target gaze angle adjustment model to obtain the target face image corresponding to the face image to be processed.

[0085] S430. Perform feature extraction on the audio information to be processed based on the audio feature extraction algorithm to obtain the acoustic features to be processed.

[0086] The audio feature extraction algorithm is an algorithm for extracting audio information features. The audio content tracking method is a method for processing the acoustic features of the audio information to be processed and obtaining the statement corresponding to the audio information to be processed, that is, a method for determining the text corresponding to the audio information to be processed.

[0087] After collecting the audio information to be processed, the audio information to be processed is subjected to acoustic feature extraction based on a preset audio feature extraction algorithm to obtain the acoustic features of the audio information to be processed.

[0088] Based on the S440, the acoustic model, and the decoder, process the acoustic features to be processed, and obtain a first statement to be determined and a first confidence coefficient corresponding to the first statement to be determined.

[0089] The acoustic model is a model for processing the extracted acoustic features to obtain the acoustic posterior probability corresponding to the acoustic features. A corresponding decoder is generated based on the content of the target text, that is, for different target texts, the decoders corresponding to different target texts are different. The first statement to be determined is the statement corresponding to the audio information to be processed after processing the acoustic features to be processed based on the acoustic model and the decoder. The first confidence coefficient is used to characterize the confidence coefficient of the first statement to be determined.

[0090] The acoustic features to be processed can be input into the acoustic model to obtain the acoustic posterior probability corresponding to the acoustic features to be processed. After inputting the acoustic posterior probability into the decoder, a first statement to be determined corresponding to the acoustic features to be processed can be obtained, and at the same time, the first confidence coefficient of the first statement to be determined can be output.

[0091] In this embodiment, the audio information to be processed can be processed based on the decoder in the audio content tracking method. First, the acoustic features to be processed are processed by the acoustic model to obtain the acoustic posterior probability corresponding to the acoustic features to be processed. Using the acoustic posterior probability as the input of the decoder, a first statement to be determined corresponding to the acoustic posterior probability and a first confidence coefficient of the first text to be determined are obtained. The first confidence coefficient is used to characterize the accuracy of the first statement to be determined. In fact, when the first confidence coefficient reaches the threshold of the preset confidence coefficient, the accuracy of the first statement to be determined is relatively high, and the first statement to be determined can be used as the statement to be matched.

[0092] S450. Based on the first confidence coefficient, determine the statement to be matched corresponding to the audio information to be processed, and determine the target statement corresponding to the target text of the statement to be matched.

[0093] When the first confidence coefficient is greater than the preset confidence coefficient threshold, the statement to be determined for the first time can be used as the statement to be matched, and at the same time, the target statement corresponding to the target text of the statement to be matched can also be determined.

[0094] Based on the above technical solution, process the acoustic features to be processed by a keyword detection system, and based on the keyword detection system in the audio content tracking method and the acoustic features to be processed, determine the second statement to be determined corresponding to the acoustic features to be processed and the second confidence coefficient corresponding to the second statement to be determined. Here, the keyword detection system matches with the target text, and when the second confidence coefficient meets the preset confidence coefficient threshold, the second statement to be determined is used as the statement to be matched.

[0095] Input the acoustic features to be processed into the keyword detection system, and the keyword detection system can output the second statement to be determined corresponding to the acoustic features to be processed and the second confidence coefficient of the second statement to be determined. When the second confidence coefficient value is higher than the preset confidence coefficient threshold, the second statement to be determined is relatively accurate, and at this time, the second statement to be determined can be used as the statement to be matched.

[0096] To improve the accuracy of the determined statement to be matched, process the acoustic features to be processed based on the decoder and the keyword detection system to determine the statement to be matched corresponding to the acoustic features to be processed.

[0097] In one embodiment, the audio content tracking method includes a keyword detection system and the decoder, processes the acoustic features to be processed based on the decoder and the keyword detection system respectively, and when obtaining a first statement to be determined and a second statement to be determined, determines the statement to be matched based on a first confidence coefficient of the first statement to be determined and a second confidence coefficient of the second statement to be determined.

[0098] Process the acoustic features to be processed based on the decoder and the keyword detection system respectively, and obtain a first statement to be determined and a second statement to be determined corresponding to the acoustic features to be processed. At the same time, obtain the confidence coefficients of the first statement to be determined and the second statement to be determined. The contents of the first statement to be determined and the second statement to be determined may be the same or different. Correspondingly, the first confidence coefficient may be the same as or different from the second confidence coefficient.

[0099] When the first statement to be determined and the second statement to be determined are the same and the first confidence coefficient and the second confidence coefficient are higher than the threshold of the preset confidence coefficient, any statement to be determined in the first statement to be determined and the second statement to be determined can be used as the statement to be matched. When the contents of the first statement to be determined and the second statement to be determined are different and the first confidence coefficient and the second confidence coefficient are higher than the threshold of the preset confidence coefficient, the statement to be determined with the higher confidence coefficient in the first statement to be determined and the second statement to be determined can be used as the statement to be matched. When the contents of the first statement to be determined and the second statement to be determined are different and the first confidence coefficient and the second confidence coefficient are lower than the threshold of the preset confidence coefficient, it means that the content currently spoken by the target user has nothing to do with the content in the target text, and it may not be necessary to determine the corresponding statement in the target text of the current audio information.

[0100] S460. Display the target statement and the target face image to the client associated with the target user respectively, or display the target statement and the target face image to the client associated with the target user simultaneously.

[0101] In this embodiment, the distinguishable display of the target statement in the target text means highlighting the target statement, or displaying the target statement in bold, or displaying other statements other than the target statement in a semi-transparent manner. Here, it includes making the transparency of a predetermined number of unbroadcast texts adjacent to the target statement lower than the transparency of other texts to be broadcast.

[0102] Highlight the target statement to prompt the user that it is the currently broadcast statement, or display the target statement in bold. Alternatively, other statements other than the target statement may be displayed in a semi-transparent manner so as not to interfere with the broadcast to the target user. Usually, in order for the target user to understand the content corresponding before and after the target statement, the transparency of a predetermined number of statements adjacent to the target statement is set to be low, so that during the broadcast process, the target user can understand the contextual meaning of the target statement, and the broadcast efficiency and usage experience of the broadcast user can be improved.

[0103] After processing the face image to be processed by the target user, the entire data stream of the target face image and the audio information related to the target face image may be sent to the client associated with the target user.

[0104] According to the technical solution of the embodiment of the present disclosure, the audio information to be processed is processed according to the audio content following method, the statement to be matched is obtained, the position of the statement to be matched in the target text is obtained, the text at the position is distinguishably displayed, intelligent following of the target text is realized, and furthermore, the technical effect of facilitating the broadcast of the target user can be achieved.

[0105] Based on the above technical solution, in order to quickly determine the target statement in the target text of the statement to be matched for broadcasting, if the broadcast statement distinguishedly displayed in the target text at the current time is included, taking the broadcast statement as the starting point, the corresponding target statement in the target text of the statement to be matched can be determined.

[0106] During the process of broadcasting to the target user, the broadcast statements and unbroadcast statements in the target text may be distinguished. For example, the broadcast statements and unbroadcast statements can be displayed with different fonts or transparencies. For example, the transparency of the broadcast text can be set high to reduce interference with the unbroadcast text. When determining the target statement at the current time, taking the last statement in the broadcast statements as the starting point, the statement that matches the statement to be matched in the unbroadcast statements is determined as the target statement, and the target statement is distinguishedly displayed.

[0107] Predetermined different distinguished display methods corresponding to different contents are set, that is, the methods for distinguishing the broadcast statements, unbroadcast statements, and target statements are different, and the technical effect of effectively prompting the broadcast user is achieved.

[0108] In the actual application process, in order to improve the determination efficiency of the target statement, starting from the broadcast statement, a predetermined number of statements after the starting point can be obtained. For example, three statements after the starting point are obtained and used as unbroadcast statements that have been registered. If there is a statement that matches the statement to be matched in the unbroadcast statements that have been registered, the statement that matches the statement to be matched can be used as the target statement. If the unbroadcast statements that have been registered do not include the statement to be matched, the target text indicates that it does not include the statement to be matched.

[0109] According to the technical solution of the embodiments of the present disclosure, when collecting the audio information to be processed, determining the acoustic features to be processed corresponding to the audio information to be processed, and inputting the acoustic features to be processed into a decoder and / or a keyword detection system corresponding to the target text, the statement to be matched corresponding to the acoustic features to be processed can be obtained, and at the same time, the sentence in the target text of the statement to be matched can be obtained, solving the problem that the teleprompter in the related art only serves to display the broadcast text and cannot effectively present it to the user, and the presentation effect is poor. During the broadcast process of the target user, collecting the audio information of the broadcast user, determining the corresponding target statement in the broadcast text of the audio information, and distinguishing and displaying it on the teleprompter, so that the teleprompter can intelligently follow the broadcast user, and the technical effect of improving the broadcast effect is achieved.

[0110] Example 5 FIG. 6 is a schematic flowchart of the data processing method provided by Embodiment 5 of the present disclosure. Based on the foregoing embodiments, first, a decoder and a keyword detection system corresponding to the target text are determined, and further, the acoustic features to be processed are processed based on the decoder and the keyword detection system. For specific embodiments, reference may be made to the description of the technical solution of the present disclosure. Here, descriptions of the same or corresponding technical terms as those in the above embodiments will not be repeated.

[0111] As shown in FIG. 6, the method includes the following steps.

[0112] S510, determine an audio content following method corresponding to the target text.

[0113] In this embodiment, to determine a decoder corresponding to the target text, it includes obtaining the target text, performing word segmentation processing on the target text to obtain at least one broadcast word corresponding to the target text, obtaining a target language model based on the at least one broadcast word, determining an interpolated language model based on the target language model and a common language model, and dynamically constructing the interpolated language model by a weighted finite state transducer to obtain a decoder corresponding to the target text.

[0114] Using various word segmentation tools such as noise word segmentation, word segmentation processing can be performed on the target text to obtain at least one broadcast word. After obtaining at least one broadcast word, a target language model corresponding to the target text can be trained, and this target language model may be a binary classification model. The common language model is a commonly used language model. An interpolated language model can be obtained based on the target language model and the common language model. The interpolated language model can determine audio unrelated to the target text spoken by the target user during the broadcast. By dynamically constructing with a weighted finite transducer, a decoder corresponding to the interpolated language model can be obtained. The decoder at this time is related to the height of the target text. Since the decoder is related to the height of the target text, a matching statement corresponding to the acoustic features to be processed based on the decoder can be effectively determined.

[0115] In this embodiment, to determine a keyword detection system corresponding to the target text, it includes splitting the target text into at least one broadcast word, determining a category corresponding to the at least one broadcast word according to a preset classification rule, and generating the keyword detection system based on the broadcast words corresponding to each category.

[0116] For example, the target text may be split into a plurality of broadcast words by a word segmentation tool. Each broadcast word may be used as one keyword. The preset classification rule may be how to classify the keywords. After determining the classification category, the broadcast words corresponding to each category are determined, and further, a keyword detection system can be generated based on the broadcast words of each category.

[0117] S520, Collect audio-video frame data related to the target user, where the audio-video frame data includes audio information to be processed and a face image to be processed.

[0118] S530. Process the face image to be processed based on the target gaze angle adjustment model to obtain a target face image corresponding to the face image to be processed.

[0119] S540. Follow-up process the audio information to be processed based on the audio content following method, and determine a target statement in the target text of the audio information to be processed.

[0120] After collecting the audio information to be processed of the target user, that is, the wav audio waveform, by the microphone array, the acoustic features to be processed can be extracted from the audio information to be processed according to the audio feature extraction method. The acoustic features to be processed can be processed by the conformer acoustic model to obtain the acoustic posterior probability. The acoustic posterior probability can be input to the decoder to obtain the first statement to be determined and the confidence coefficient corresponding to the first statement to be determined. At the same time, the acoustic features to be processed can be input to the keyword detection system to obtain the second statement to be determined corresponding to the acoustic features to be processed and the confidence coefficient corresponding to the second statement to be determined. The statements to be determined corresponding to the two confidence coefficients are fused to determine the statement to be matched. For example, the statement to be determined with a high confidence coefficient can be used as the statement to be matched.

[0121] In this embodiment, to determine the target statement corresponding to the target text of the statement to be matched, starting from the last statement of the currently broadcast text in the target text, it is determined whether the next text matches the statement to be matched. If the next text matches the statement to be matched, it is determined that the next text is the target statement. If the next text does not match the statement to be matched, it is determined whether the next statement of the next text matches the statement to be matched. If the next statement of the next text matches the statement to be matched, it is determined that the statement is the target statement. If there is no target statement that matches the statement to be matched, it can be left unprocessed because the current utterance of the target user has nothing to do with the target text.

[0122] S550. Display the target statement and the target face image to the client associated with the target user respectively, or display the target statement and the target face image to the client associated with the target user simultaneously.

[0123] The target statement can be distinguished from other statements in the target text.

[0124] Based on the above technical solution, during the determination process of the target statement, determining the actual audio duration corresponding to the target statement, adjusting the predicted audio duration corresponding to the unmatched statement based on the actual audio duration and the unmatched statements in the target text, and displaying the predicted audio duration to the target client to prompt the target user are included.

[0125] The actual audio duration refers to the duration used by the target user to say the target statement, for example, 2S. The predicted audio duration refers to the duration used by the target user to broadcast subsequent unbroadcast statements. The predicted broadcast duration is a dynamically adjusted duration, which is mainly determined according to the speech rate of the target user. The speech rate is determined by determining the duration used by each word based on the actual audio duration used by the target statement and the number of words corresponding to the target statement. Based on the duration used by each word and the total number of words in the subsequent unbroadcast statement, the duration required for broadcasting the subsequent unbroadcast statement can be determined. The unbroadcast statement can be used as an unmatched statement.

[0126] In the actual application process, in order to achieve the effect of timely prompting the user, the predicted audio duration may be displayed to the target client corresponding to the target user. At the same time, the target user may adjust the speech rate of the broadcast text based on the predicted audio duration so that the broadcast duration matches the preset duration, that is, complete the broadcast of the content of the target text within the limited duration.

[0127] For the target user, the duration used for broadcasting each statement is different, and the corresponding duration used for broadcasting each word is also different. During the broadcasting process of the target user, the predicted audio duration may be dynamically adjusted according to the current broadcast duration of each word.

[0128] Based on the above technical solution, when the method receives the target text, it tags sentence delimiters for the target text, displays the sentence delimiter tagging identifier to the client, and the user can view the target text based on the sentence delimiter tagging identifier.

[0129] With the popularization of video shooting, not all video creators have the opportunity to receive professional training related to broadcasting or emceeing. Therefore, as a teleprompter or application, it will be more versatile if video creators with no broadcasting foundation can create higher-quality voice broadcast videos. For general users who do not have the ability to analyze specialized word breaks, they cannot determine at what timing to break the input target text or what emotions to process each broadcast text with. Therefore, in addition to the above technical effects, this technical solution further includes a sentence segmentation tagging model. After the target text is uploaded, the target text is segmented and tagged, and the sentence segmentation tagging result is displayed on the target terminal used by the target user, enabling the user to broadcast the target text based on the sentence segmentation tagging and improving the professionalism of the broadcast target text.

[0130] The sentence segmentation tagging identifier may be displayed as " / ". For example, when a long pause is required, it may be displayed as "--", and when a short pause is required, it may be displayed as "-". When two words need to be read together, it may be displayed as "()". At the same time, the sentence segmentation tagging identifier may be displayed in the target text and the target text may be displayed on the target terminal.

[0131] After uploading the target text to the server corresponding to the mobile terminal, the target text is segmented and tagged. When the target user broadcasts, it is possible to quickly judge how to broadcast according to the sentence segmentation tagging, solving the problem of low efficiency that requires manual judgment of how to segment and tag the content of the broadcast text in the related technology.

[0132] In this embodiment, the sentence segmentation tagging of the target text is realized by a pre-trained sentence segmentation tagging model.

[0133] Based on the above technical solution, the method further includes tagging the broadcast statements in the target text with broadcast identifiers and determining the broadcast statements and unbroadcast statements in the target text based on the broadcast identifiers.

[0134] During the user's broadcast process, different colors are attached to the broadcast statements and unbroadcast statements in the target text, so that the target user can clearly distinguish the broadcast statements and unbroadcast statements during the broadcast. At the same time, the target content is distinguished and displayed, which can clearly prompt the target user about the content currently being read and the content to be read next, and can solve the problem of deviation during the text broadcast for the target user.

[0135] Based on the above technical solution, when the method receives the uploaded target text, it performs emotion tagging on each statement in the target text based on a pre-learned emotion tagging model, and the user can broadcast the target text based on the emotion tagging. That is, when receiving the target text, emotion tagging is performed on each statement in the target text, and an emotion tagging identifier is displayed to the client, so that the user can view the target text based on the emotion tagging identifier.

[0136] For general users, in addition to not knowing at which position in the target text to pause, usually they also do not know what kind of emotional color to use to broadcast the text. At this time, after receiving the uploaded target text, the target text can be pre-processed. For example, perform emotion color analysis on each statement in the target text, perform expression tagging on each analyzed statement, and solve the problem of rigid browsing by having the user broadcast the content in the target text based on the emotion color identifier.

[0137] According to the technical solution of the embodiments of the present disclosure, word segmentation processing is performed on the uploaded target text to obtain a decoder and a keyword detection system corresponding to the target text, the acoustic features extracted based on the decoder and the keyword detection system are processed, it is determined whether the content of the target text is what the user is speaking, and it is distinguished and displayed, and while achieving the effect of intelligently following the customized user with the distinguished and displayed text, in order to enhance the broadcast effect, the target text is segmented into sentences and tagged, and emotion tagged, so that the user can broadcast the text based on the displayed sentence segmentation tagging and emotion tagging, and a three-dimensional broadcast of the broadcast text can be realized.

[0138] Example 6 FIG. 7 is a schematic structural diagram of an intelligent teleprompter provided by Example 6 of the present application. The intelligent teleprompter realizes the above data processing method. For example, as shown in FIG. 7, the intelligent teleprompter is configured at the application level, and the prompting method is deployed in the application. The application may be an application installed on a mobile terminal and capable of implementing live broadcast interaction. The target text may be displayed on the display interface, and the target statement corresponding to the audio information of the target user may be distinguished and displayed. At the same time, based on the algorithm for analyzing and processing the audio information, the target text can be intelligently segmented into words, emotion tagged, intelligently segmented into sentences and tagged, and semantic analyzed, and intelligent following, read tag, and duration prediction of unread statements can be realized. The implementation manner may refer to the above embodiments. At the same time, the face image of the target user is further displayed on the display interface, and the model is adjusted based on the target line-of-sight angle in the algorithm to focus on the line of sight of the target user, and a target face image with the line-of-sight angle being the target angle can be obtained, and the target face image can be transmitted to the live broadcast viewing client.

[0139] The application scenario, technical solution, and adopted algorithm of this embodiment are shown in FIG. 8. The application scenario may be any scenario for interaction, such as an online education scenario, a live broadcast scenario in the field of e-commerce, a content video recording scenario, or an online lecture scenario, etc. The implementation solution is intelligent following, read tagging, time estimation, expression tagging, eye focus, and intelligent sentence segmentation, and the algorithm adopted for the realization of the above functions may be a model training algorithm and an audio intelligent analysis algorithm.

[0140] Example 7 FIG. 9 is a schematic structural diagram of a data processing device provided by Example 7 of the present disclosure, which can execute the data processing method provided by any embodiment of the present disclosure and has a functional module and effect corresponding to the execution method. As shown in FIG. 9, the device includes an audio-video frame data collection module 610, a face image processing module 620, a target statement determination module 630, and a display module 640.

[0141] The audio-video frame data collection module 610 is configured to collect audio-video frame data related to a target user. Here, the audio-video frame data includes audio information to be processed and a face image to be processed. The face image processing module 620 is configured to process the face image to be processed based on a target gaze angle adjustment model to obtain a target face image corresponding to the face image to be processed. The target statement determination module 630 is configured to perform follow-up processing on the audio information to be processed based on an audio content following method and determine a target statement in the target text of the audio information to be processed. The display module 640 is configured to display the target statement and the target face image on a client related to the target user respectively, or to display the target statement and the target face image on a client related to the target user simultaneously.

[0142] Based on the above technical solution, before collecting the audio-video frame data related to the target user, the audio-video frame data collection module 610 is further configured to upload the target text to enable the target user to interact based on the target text.

[0143] Based on the above technical solution, when it is detected that a preset event is triggered, the audio-video frame data collection module 610 is further configured to collect the audio information to be processed by the target user based on the audio collection device and collect the face image to be processed by the target user based on the photographing device.

[0144] Based on the above technical solution, the face image processing module 620 is further configured to input the face image to be processed into the target line-of-sight angle adjustment model to obtain the target face image, where the line-of-sight angle of the target user in the target face image coincides with the target line-of-sight angle.

[0145] Based on the above technical solution, the device further includes a target line-of-sight angle adjustment model training module, which A sample set acquisition unit configured to acquire a training sample set, where the training sample set includes a plurality of training samples, each training sample includes a target line-of-sight angle image and a non-target line-of-sight angle image, the training samples are determined based on a target sample generation model obtained by pre-training, for each training sample, input the non-target line-of-sight angle image in the current training sample into a line-of-sight angle adjustment model to be trained, and a target line-of-sight angle adjustment model training unit configured to obtain an actual output image corresponding to the current training sample, and based on the actual output image and the target line-of-sight angle image of the current training sample, determine a loss value, and based on the loss value and a preset loss function of the line-of-sight angle adjustment model to be trained, adjust the model parameters of the line-of-sight angle adjustment model to be trained, converge the preset loss function as a training target, and obtain the target line-of-sight angle adjustment model.

[0146] Based on the above technical solution, the apparatus further comprises a sample generation model training module configured to obtain the target sample generation model by training, where the target sample generation model comprises a target line-of-sight angle image generation sub-model and a non-target line-of-sight angle image generation sub-model.

[0147] Based on the above technical solution, the sample generation model training module An error determination unit configured to process a pre-collected Gaussian distribution vector based on a generator in a non-target line-of-sight angle image generation sub-model to be trained to obtain an image to be compared, and to process an original non-frontal view sample image and the image to be compared based on a discriminator in the non-target line-of-sight angle image generation sub-model to be trained to obtain an error, where the original non-frontal view sample image is a pre-collected image, and a parameter correction unit configured to correct model parameters in the non-target line-of-sight angle image generation sub-model to be trained based on the error and a loss function in the non-target line-of-sight angle image generation sub-model to be trained, and a sub-model generation unit configured to converge the loss function as a training target to obtain the non-target line-of-sight angle image generation sub-model and to generate a non-target line-of-sight angle image in the training sample based on the non-target line-of-sight angle image generation sub-model.

[0148] Based on the above technical solution, the sample generation model training module further obtains model parameters in the non-target line-of-sight angle image generation sub-model, multiplies the model parameters by a target line-of-sight angle image generation sub-model to be trained, trains the target line-of-sight angle image generation sub-model to be trained based on a pre-collected Gaussian distribution vector and an original frontal view sample image to obtain a target line-of-sight angle image generation sub-model, and is configured to generate a target line-of-sight angle image in the training sample based on the target line-of-sight angle image generation sub-model.

[0149] Based on the above technical solution, the audio content following method includes an audio feature extraction algorithm and a decoder, and the target statement determination module 630 Perform feature extraction on the audio information to be processed based on the audio feature extraction algorithm to obtain the acoustic features to be processed, process the acoustic features to be processed based on an acoustic model to obtain the acoustic posterior probability corresponding to the acoustic features to be processed, and based on the acoustic posterior probability and the decoder, determine a first statement to be determined and a first confidence coefficient corresponding to the first statement to be determined. Here, the decoder is determined based on an interpolation language model corresponding to the target text, and the interpolation language model is determined based on a target language model and a common language model corresponding to the target text. When the first confidence coefficient meets the threshold of the preset confidence coefficient, use the first statement to be determined as the statement to be matched, and configure to determine the target statement based on the statement to be matched.

[0150] Based on the above technical solution, the audio content tracking method includes a keyword detection system, and the target statement determination module 630 Process the acoustic features to be processed of the audio information to be processed based on the keyword detection system, and determine a second statement to be determined corresponding to the acoustic features to be processed and a second confidence coefficient of the second statement to be determined. When the second confidence coefficient meets the threshold of the preset confidence coefficient, use the second statement to be determined as the statement to be matched, and configure to determine the target statement based on the statement to be matched.

[0151] Based on the above technical solution, the target statement determination module 630 The audio content following method includes a keyword detection system and the decoder, and based on the decoder and the keyword detection system, processes the acoustic features to be processed respectively. When obtaining a first statement to be determined and a second statement to be determined, based on a first confidence coefficient of the first statement to be determined and a second confidence coefficient of the second statement to be determined, determines a statement to be matched, and is configured to determine the target statement based on the statement to be matched.

[0152] Based on the above technical solution, the display module 640 is configured to separately display the target statement to the first client as the target text, and display a target audio-video frame corresponding to the target face image to the second client.

[0153] Based on the above technical solution, the apparatus further includes a decoder generation module, which determines a decoder corresponding to the target text. Determining the decoder corresponding to the target text includes obtaining the target text, performing word segmentation processing on the target text to obtain at least one broadcast word corresponding to the target text, obtaining a target language model based on the at least one broadcast word, determining an interpolated language model based on the target language model and a common language model, and dynamically constructing the interpolated language model by a weighted finite state transducer to obtain a decoder corresponding to the target text.

[0154] Based on the above technical solution, the apparatus further includes a keyword detection system generation module, which splits the target text into at least one broadcast word, determines a category corresponding to the at least one broadcast word according to a pre-determined classification rule, and is configured to generate the keyword detection system based on the broadcast words corresponding to each category.

[0155] Based on the above technical solution, the display module 640 includes a discrimination display unit, which is configured to highlight the target statement, or display the target statement in bold, or display other statements other than the target statement in a semi-transparent manner, where the transparency of a predetermined number of unbroadcast texts adjacent to the target statement is lower than the transparency of other texts to be broadcast.

[0156] Based on the above technical solution, the display module 640 includes a predicted duration unit, which determines the actual audio duration corresponding to the audio information to be processed, adjusts the predicted audio duration corresponding to the unviewed statement based on the actual audio duration and the unviewed statements in the target text, and displays the predicted audio duration to the target client to which the target user belongs to present the target user.

[0157] Based on the above technical solution, the device further includes a sentence segmentation tagging module, which when receiving the target text, tags sentence segmentation for the target text, and is configured to display the sentence segmentation tagging identifier to the client for the user to view the target text based on the sentence segmentation tagging identifier.

[0158] Based on the above technical solution, the device further includes an emotion tagging module, which when receiving the target text, tags emotion for each statement in the target text, and is configured to display the emotion tagging identifier to the client for the user to view the target text based on the emotion tagging identifier.

[0159] According to the technical solution of the embodiments of the present disclosure, audio information to be processed by a target user and a face image to be processed are collected, the audio information to be processed and the face image to be processed are processed, a target face image and a position in the target text of a target statement corresponding to the audio information to be processed can be obtained, solving the problem that the occupied area of a teleprompter in the related art is large, the use is complicated, and the use of the teleprompter is inconvenient, and achieving the technical effect that it can be efficiently broadcast only by a mobile terminal.

[0160] The multiple units and modules included in the above device are divided according to functional logic. As long as the corresponding functions can be realized, the division is not limited to the above, and the names of the multiple functional units are for distinguishing from each other and do not limit the protection scope of the embodiments of the present disclosure.

[0161] Embodiment 7 FIG. 10 is a schematic structural diagram of an electronic device provided by Embodiment 8 of the present disclosure. Hereinafter, referring to FIG. 10, it is a schematic structural diagram of an electronic device (for example, the terminal device or the server in FIG. 10) 700 suitable for the implementation of the embodiments of the present disclosure. The terminal device 700 in the embodiments of the present disclosure includes, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions, TVs), desktop computers, etc. The electronic device 700 shown in FIG. 10 is merely an example and does not limit the functions and usage ranges of the embodiments of the present disclosure in any way.

[0162] As shown in FIG. 10, the electronic device 700 includes a processing device (e.g., a central processing unit, a graphics processor, etc.) 701 that can execute various appropriate operations and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data necessary for the operation of the electronic device 700 are further stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0163] Normally, the I / O interface 705 is connected to an input device 706 such as a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, an output device 707 such as a liquid crystal display (LCD), a speaker, a vibrator, a storage device 708 such as a magnetic tape, a hard disk, and a communication device 709. The communication device 709 can enable the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although FIG. 10 illustrates the electronic device 700 equipped with various devices, it should be understood that it is not necessary to implement or include all the illustrated devices. Alternatively, more or fewer devices may be implemented or included.

[0164] According to the embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments of the present disclosure provide a computer program product including a computer program carried on a non-transitory computer-readable medium, where the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via a communication device 709, or may be installed from a storage device 708, or may be installed from a ROM 702. When the computer program is executed by a processing device 701, the above functions defined in the method of the embodiments of the present disclosure are realized.

[0165] The names of messages or information that interact between multiple devices in the embodiments of the present disclosure are used only for the purpose of description and are not intended to limit the scope of those messages or information.

[0166] The electronic device provided by the embodiments of the present disclosure and the data processing method provided by the above embodiments belong to the same concept. Technical details not comprehensively described in this embodiment may be referred to the above embodiments, and this embodiment has the same effects as the above embodiments.

[0167] Embodiment 8 The embodiments of the present disclosure provide a computer storage medium storing a computer program, and when the program is executed by a processor, the data processing method provided by the above embodiments is realized.

[0168] The computer-readable media described in this disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof, but is not limited thereto. Examples of computer-readable storage media include electrical connections having one or more conductors, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above, but are not limited thereto. In this disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program, and this program may be used by, or in combination with, an instruction execution system, apparatus, or device. In this disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier, and carry computer-readable program code. Such a propagated data signal includes, but is not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium that transmits, propagates, or transfers a program used by, or in combination with, an instruction execution system, apparatus, or device. The program code included in the computer-readable medium may be transferred by any suitable medium, such as a wire, optical fiber cable, radio frequency (RF), or any suitable combination of the above, but is not limited thereto.

[0169] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol, such as the HyperText Transfer Protocol (HTTP), and may be interconnected with digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), internetworks (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), and any currently known or future-developed network.

[0170] The computer-readable medium may be included in the electronic device or may be separate and not incorporated into the electronic device.

[0171] The computer-readable storage medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to collect audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed and face images to be processed, process the face images to be processed based on a target line-of-sight angle adjustment model to obtain target face images corresponding to the face images to be processed, perform follow-up processing on the audio information to be processed based on an audio content following method, determine relevant target statements in the target text of the audio information to be processed, and display the target statements and the target face images on a client related to the target user respectively, or cause the target statements and the target face images to be simultaneously displayed on a client related to the target user.

[0172] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof, including, but not limited to, object-oriented programming languages (such as Java, Smalltalk, C++), and conventional procedural programming languages (such as the "C" language or similar programming languages). The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, such as a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functionality, and operation of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the boxes may occur in a different order than shown in the accompanying drawings. For example, two boxes shown in succession may in fact be executed substantially in parallel, or depending on the related functionality, may be executed in the reverse order. Also, each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0174] The units described in the embodiments of the present disclosure may be implemented by software or by hardware. Here, the name of the unit does not constitute a limitation of the unit itself in a given situation. For example, the first acquisition unit may also be described as "a unit for acquiring at least two Internet protocol addresses".

[0175] The functions described above in this specification may be executed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include Field-Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System On Chip (SOC), Complex programmable logic device (CPLD), and the like. In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program used by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable medium. The machine-readable medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of the machine-readable storage medium include electrical connections based on one or more wires, portable device disks, hard disks, RAM, ROM, EPROM, or flash memory, optical fibers, CD-ROM, optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0176] According to one or more embodiments of the present disclosure, [Example 1] provides a data processing method, the method comprising: collecting audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed and a face image to be processed; processing the face image to be processed based on a target gaze angle adjustment model to obtain a target face image corresponding to the face image to be processed; Performing follow-up processing on the audio information to be processed based on the audio content follow-up method, and determining a target statement in the target text of the audio information to be processed. Including displaying the target statement and the target face image to the client associated with the target user respectively, or simultaneously displaying the target statement and the target face image to the client associated with the target user.

[0177] According to one or more embodiments of the present disclosure, [Example 2] provides a data processing method. Here, before collecting the audio-video frame data associated with the target user, Further including receiving the uploaded target text in order to enable the target user to interact based on the target text.

[0178] According to one or more embodiments of the present disclosure, [Example 3] provides a data processing method. Here, collecting the audio-video frame data associated with the target user includes When it is detected that a preset event is triggered, collecting the audio information to be processed of the target user based on an audio collection device, and collecting the face image to be processed of the target user based on a photographing device.

[0179] According to one or more embodiments of the present disclosure, [Example 4] provides a data processing method. Here, performing focusing processing on the face image information based on the target line-of-sight angle adjustment model obtained by the pre-learning, and obtaining a target face image corresponding to the face image to be processed includes Inputting the face image to be processed into the target line-of-sight angle adjustment model to obtain the target face image, where the line-of-sight angle of the target user in the target face image is consistent with the target line-of-sight angle.

[0180] According to one or more embodiments of the present disclosure, [Example 5] provides a data processing method, where obtaining a training sample set, where the training sample set includes a plurality of training samples, each training sample includes a target line-of-sight angle image and a non-target line-of-sight angle image, and the training sample is determined based on a target sample generation model obtained by pre-training. For each training sample, input the non-target line-of-sight angle image in the current training sample into the line-of-sight angle adjustment model to be trained, and obtain the actual output image corresponding to the current training sample. Based on the actual output image and the target line-of-sight angle image of the current training sample, determine a loss value, and based on the loss value and a preset loss function of the line-of-sight angle adjustment model to be trained, adjust the model parameters of the line-of-sight angle adjustment model to be trained. Converge the preset loss function as a training target to obtain the target line-of-sight angle adjustment model.

[0181] According to one or more embodiments of the present disclosure, [Example 6] provides a data processing method, where obtaining the non-target line-of-sight angle image generation sub-model in the target sample generation model by training includes: Processing a pre-collected Gaussian distribution vector based on the generator in the non-target line-of-sight angle image generation sub-model to be trained to obtain an image to be compared. Processing the original non-frontal view sample image and the image to be compared based on the discriminator in the non-target line-of-sight angle image generation sub-model to be trained to obtain an error, where the original non-frontal view sample image is a pre-collected image. Based on the error and a loss function in the non-target line-of-sight angle image generation sub-model to be trained, correct the model parameters in the non-target line-of-sight angle image generation sub-model to be trained. Converge the loss function as a training target to obtain the non-target line-of-sight angle image generation sub-model, and generate the non-target line-of-sight angle image in the training sample based on the non-target line-of-sight angle image generation sub-model.

[0182] According to one or more embodiments of the present disclosure, [Example 7] provides a data processing method, wherein the training for obtaining the target line-of-sight angle image generation sub-model in the target sample generation model is acquiring the model parameters in the non-target line-of-sight angle image generation sub-model and multiplying the model parameters by the target line-of-sight angle image generation sub-model to be trained training the target line-of-sight angle image generation sub-model to be trained based on the pre-collected Gaussian distribution vector and the original front view sample image to obtain a target line-of-sight angle image generation sub-model, and generating a target line-of-sight angle image in the training sample based on the target line-of-sight angle image generation sub-model.

[0183] According to one or more embodiments of the present disclosure, [Example 8] provides a data processing method, wherein the audio content following method includes an audio feature extraction algorithm and a decoder corresponding to the target text, and following the audio information to be processed based on the audio content following method, and determining the relevant target statement in the target text of the audio information to be processed is performing feature extraction on the audio information to be processed based on the audio feature extraction algorithm to obtain the acoustic features to be processed processing the acoustic features to be processed based on an acoustic model to obtain the acoustic posterior probability corresponding to the acoustic features to be processed determining a first statement to be determined and a first confidence coefficient corresponding to the first statement to be determined based on the acoustic posterior probability and the decoder, wherein the decoder is determined based on an interpolation language model corresponding to the target text, and the interpolation language model is determined based on a target language model and a common language model corresponding to the target text when the first confidence coefficient meets the threshold of the preset confidence coefficient, using the first statement to be determined as the statement to be matched, and determining the target statement based on the statement to be matched.

[0184] According to one or more embodiments of the present disclosure, [Example 9] provides a data processing method, where the audio content tracking method includes a keyword detection system, and tracking and processing the audio information to be processed based on the audio content tracking method, and determining a related target statement in the target text of the audio information to be processed, processing the acoustic features to be processed of the audio information to be processed based on the keyword detection system, and determining a second statement to be determined corresponding to the acoustic features to be processed and a second confidence coefficient of the second statement to be determined, when the second confidence coefficient meets the threshold of the preset confidence coefficient, using the second statement to be determined as the statement to be matched, and determining the target statement based on the statement to be matched.

[0185] According to one or more embodiments of the present disclosure, [Example 10] provides a data processing method, where the audio content tracking method includes a keyword detection system and the decoder, and when the decoder and the keyword detection system are respectively used to process the acoustic features to be processed to obtain a first statement to be determined and a second statement to be determined, determining the statement to be matched based on the first confidence coefficient of the first statement to be determined and the second confidence coefficient of the second statement to be determined, and determining the target statement based on the statement to be matched. The audio content tracking method includes a keyword detection system and the decoder, and when the decoder and the keyword detection system are respectively used to process the acoustic features to be processed to obtain a first statement to be determined and a second statement to be determined, determining the statement to be matched based on the first confidence coefficient of the first statement to be determined and the second confidence coefficient of the second statement to be determined, and determining the target statement based on the statement to be matched.

[0186] According to one or more embodiments of the present disclosure, [Example 11] provides a data processing method, wherein displaying the target statement and the target face image to the client associated with the target user respectively, or displaying the target statement and the target face image to the client associated with the target user simultaneously is including separately displaying the target statement as target text on a first client and displaying a target audio-video frame corresponding to the target face image on a second client.

[0187] According to one or more embodiments of the present disclosure, [Example 12] provides a data processing method, wherein determining a decoder corresponding to the target text Determining the decoder corresponding to the target text includes obtaining the target text, performing word segmentation processing on the target text to obtain at least one broadcast word corresponding to the target text, obtaining a target language model based on the at least one broadcast word, determining an interpolated language model based on the target language model and a common language model, and dynamically constructing the interpolated language model by a weighted finite state transducer to obtain a decoder corresponding to the target text.

[0188] According to one or more embodiments of the present disclosure, [Example 13] provides a data processing method, wherein determining a keyword detection system corresponding to the target text includes segmenting the target broadcast text into at least one broadcast word, determining a category corresponding to the at least one broadcast word according to a pre-determined classification rule, and generating the keyword detection system based on the broadcast words corresponding to each category.

[0189] According to one or more embodiments of the present disclosure, [Example 14] provides a data processing method, where determining the corresponding target statement in the target text of the statement to be matched includes when the target text at the current time includes a viewed statement that is distinctively displayed, determining the corresponding target statement in the target text of the statement to be matched with the viewed statement as the starting point.

[0190] According to one or more embodiments of the present disclosure, [Example 15] provides a data processing method, where determining the corresponding target statement in the target text of the statement to be matched with the broadcast statement as the starting point includes determining a predetermined number of registered but unviewed statements after the starting point with the broadcast statement as the starting point, when there is a statement in the registered but unviewed statements that matches the statement to be matched, using the statement that matches the statement to be matched as the target statement.

[0191] According to one or more embodiments of the present disclosure, [Example 16] provides a data processing method, where distinctively displaying the target statement in the target text includes highlighting the target statement, or displaying the target statement in bold, or displaying other statements other than the target statement in semi - transparency, where the transparency of a predetermined number of unbroadcast texts adjacent to the target statement is lower than the transparency of other texts to be broadcast.

[0192] According to one or more embodiments of the present disclosure, [Example 17] provides a data processing method, where after determining the target statement, Determining an actual audio duration corresponding to the audio information to be processed; Adjusting a predicted audio duration corresponding to the unviewed statement based on the actual audio duration and the unviewed statements in the target text; Including presenting the target user by displaying the predicted audio duration to a target client to which the target user belongs.

[0193] According to one or more embodiments of the present disclosure, [Example 18] provides a data processing method, wherein when the target text is received, sentence delimiters are tagged to the target text, and sentence delimiter identifiers are displayed to the client for the user to view the target text based on the sentence delimiter tagging identifiers.

[0194] According to one or more embodiments of the present disclosure, [Example 19] provides a data processing method, wherein when the target text is received, each statement in the target text is emotion-tagged, and emotion tagging identifiers are displayed to the client for the user to view the target text based on the emotion tagging identifiers.

[0195] According to one or more embodiments of the present disclosure, [Example 20] provides a data processing apparatus, the apparatus comprising: An audio-video frame data collection module configured to collect audio-video frame data related to a target user, wherein the audio-video frame data includes audio information to be processed and a face image to be processed; A face image processing module configured to process the face image to be processed based on a target line-of-sight angle adjustment model to obtain a target face image corresponding to the face image to be processed; A target statement determination module configured to perform follow-up processing on the audio information to be processed based on an audio content following method and determine related target statements in the target text of the audio information to be processed; A display module configured to display the target statement and the target face image to a client associated with the target user respectively, or to display the target statement and the target face image to the client associated with the target user simultaneously.

[0196] Furthermore, although a plurality of operations are depicted using a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although details of multiple implementations are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Specific features described in the context of a single embodiment may also be implemented in combination within a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually, or in any suitable sub-combination, in multiple embodiments.

Claims

1. Collecting audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed by the target user and a face image to be processed by the target user, and processing the face image to be processed based on a target line-of-sight angle adjustment model to obtain a target face image with the line-of-sight angle of the target user in the face image to be processed adjusted to a target line-of-sight angle, and performing follow-up processing on the audio information to be processed based on an audio content following method to determine a target statement in target text associated with the audio information to be processed, and displaying the target statement on at least a client used by the target user, and displaying the target face image on at least a client used by an audience user, the data processing method comprising.

2. Before collecting the audio-video frame data related to the target user, receiving the uploaded target text to enable the target user to interact based on the target text, further comprising the method according to claim 1.

3. Collecting the audio-video frame data related to the target user comprises, when it is detected that a preset event is triggered, collecting the audio information to be processed by the target user based on an audio collection device, and collecting the face image to be processed by the target user based on a photographing device, the method according to claim 1.

4. processing the face image to be processed based on the target line-of-sight angle adjustment model to obtain a target face image corresponding to the face image to be processed, and inputting the face image to be processed into the target line-of-sight angle adjustment model to obtain the target face image, where the line-of-sight angle of the target user in the target face image is consistent with the target line-of-sight angle, the method according to claim 1.

5. Obtaining a training sample set, where the training sample set includes a plurality of training samples, each training sample includes a target line-of-sight angle image and a non-target line-of-sight angle image, and the training samples are determined based on a target sample generation model obtained by pre-training, and For each training sample, input the non-target line-of-sight angle image in the current training sample into the line-of-sight angle adjustment model to be trained, and obtain the actual output image corresponding to the current training sample. Based on the actual output image and the target line-of-sight angle image of the current training sample, determine a loss value, and based on the loss value and the preset loss function of the line-of-sight angle adjustment model to be trained, adjust the model parameters of the line-of-sight angle adjustment model to be trained. Converge the preset loss function of the line-of-sight angle adjustment model to be trained as a training target to obtain the target line-of-sight angle adjustment model. The method according to claim 1 includes the above steps.

6. Further include obtaining, by training, the non-target line-of-sight angle image generation sub-model in the target sample generation model. The training for obtaining the non-target line-of-sight angle image generation sub-model in the target sample generation model is as follows: Process the pre-collected Gaussian distribution vector based on the generator in the non-target line-of-sight angle image generation sub-model to be trained to obtain an image to be compared. Process the original non-frontal view sample image and the image to be compared based on the discriminator in the non-target line-of-sight angle image generation sub-model to be trained to obtain an error. Here, the original non-frontal view sample image is a pre-collected image. Based on the error and the loss function in the non-target line-of-sight angle image generation sub-model to be trained, correct the model parameters in the non-target line-of-sight angle image generation sub-model to be trained. Converge the loss function in the non-target line-of-sight angle image generation sub-model to be trained as a training target to obtain the non-target line-of-sight angle image generation sub-model, and generate the non-target line-of-sight angle image in the training sample based on the non-target line-of-sight angle image generation sub-model. The method according to claim 5 includes the above steps.

7. Further include obtaining, by training, the target line-of-sight angle image generation sub-model in the target sample generation model. The training for obtaining the target line-of-sight angle image generation sub-model in the target sample generation model is as follows: Obtain the model parameters in the non-target line-of-sight angle image generation sub-model, and multiply the model parameters by the target line-of-sight angle image generation sub-model to be trained. Training the target gaze angle image generation sub-model to be trained based on the pre-collected Gaussian distribution vectors and the original front-view sample images to obtain the target gaze angle image generation sub-model, and generating the target gaze angle images in the training samples based on the target gaze angle image generation sub-model, the method according to claim 6, comprising:

8. The audio content tracking method includes an audio feature extraction algorithm and a decoder corresponding to the target text, Tracking the audio information to be processed based on the audio content tracking method, and determining the target statement in the target text associated with the audio information to be processed, Performing feature extraction on the audio information to be processed based on the audio feature extraction algorithm to obtain the acoustic features to be processed, Processing the acoustic features to be processed based on an acoustic model to obtain the acoustic posterior probability corresponding to the acoustic features to be processed, Determining a first statement to be determined and a first confidence coefficient corresponding to the first statement to be determined based on the acoustic posterior probability and the decoder, where the decoder is determined based on an interpolation language model corresponding to the target text, and the interpolation language model is determined based on a target language model and a common language model corresponding to the target text, When the first confidence coefficient meets a preset confidence coefficient threshold, using the first statement to be determined as the statement to be matched, and determining the target statement based on the statement to be matched, the method according to claim 1, comprising:

9. The audio content tracking method includes a keyword detection system, Tracking the audio information to be processed based on the audio content tracking method, and determining the target statement in the target text associated with the audio information to be processed, Processing the acoustic features to be processed of the audio information to be processed based on the keyword detection system to determine a second statement to be determined and a second confidence coefficient of the second statement to be determined corresponding to the acoustic features to be processed, When the second confidence coefficient satisfies a preset confidence coefficient threshold value, regarding the second statement to be determined as a statement to be matched, and determining the target statement based on the statement to be matched, the method according to claim 1 includes.

10. Tracking and processing the audio information to be processed based on the audio content tracking method, and determining a target statement in the target text associated with the audio information to be processed is The audio content tracking method includes a keyword detection system and a decoder. When processing the acoustic features to be processed of the acoustic features to be processed based on the decoder and the keyword detection system respectively to obtain a first statement to be determined and a second statement to be determined, based on the first confidence coefficient of the first statement to be determined and the second confidence coefficient of the second statement to be determined, determining a statement to be matched, and determining the target statement based on the statement to be matched, the method according to claim 1 includes.

11. Displaying the target statement and the target face image on the client associated with the target user respectively, or simultaneously displaying the target statement and the target face image on the client associated with the target user is The method according to claim 1 includes separately displaying the target statement in the target text on a first client and displaying a target audio / video frame corresponding to the target face image on a second client.

12. During the process of determining the target statement Determining the actual audio duration corresponding to the audio information to be processed Adjusting the predicted audio duration corresponding to the unviewed statement based on the actual audio duration and the unviewed statements in the target text Further including displaying the predicted audio duration on the target client to which the target user belongs and presenting the target user, the method according to claim 1.

13. The method according to claim 1, further comprising: when receiving the target text, tagging sentence delimiters to the target text, displaying a sentence delimiter tagging identifier to the client, and enabling the target user to view the target text based on the sentence delimiter tagging identifier.

14. The method according to claim 1, further comprising: when receiving the target text, performing emotion tagging on each statement in the target text, displaying an emotion tagging identifier to the client, and enabling the target user to view the target text based on the emotion tagging identifier.

15. An audio-video frame data collection module configured to collect audio-video frame data related to a target user, where the audio-video frame data includes audio information to be processed by the target user and a face image to be processed by the target user. A face image processing module configured to process the face image to be processed based on a target gaze angle adjustment model and obtain a target face image with the gaze angle of the target user in the face image to be processed adjusted to a target gaze angle. A target statement determination module configured to perform follow-up processing on the audio information to be processed based on an audio content following method and determine a target statement in target text associated with the audio information to be processed. A data processing apparatus comprising: a display module configured to display the target statement to at least a client used by the target user and display the target face image to at least a client used by a viewing user.

16. An electronic device comprising at least one processor and a storage device storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor is caused to execute the data processing method according to any one of claims 1 to 14.

17. A storage medium containing computer-executable instructions, wherein when the computer-executable instructions are executed by a computer processor, the data processing method according to any one of claims 1 to 14 is executed. ​

Citation Information

Patent Citations

  • Platform system

    JP1998290717A

  • Character information display controller, and program

    JP2011013542A

  • Image processing apparatus, computer program, video call system, and image processing method

    JP2019201360A

  • Image processing device, image processing method, and program

    WO2020110811A1