Data processing method, apparatus, device, and medium

By matching and highlighting prompts corresponding to the user's voice in real time during short video recording, the problem of inexperienced creators forgetting their lines is solved, thus improving the quality of video recording.

CN114911448BActive Publication Date: 2025-11-07TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110179007.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-08
Publication Date
2025-11-07
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

During the recording of short videos, inexperienced multimedia creators are prone to problems such as forgetting lines or mispositioning, which leads to a decline in video quality.

Method used

By capturing user voice during video recording, matching and highlighting prompt text associated with the voice, and using a backend server for voice-to-text matching, the text is displayed in real time on the recording page, ensuring synchronization between text and voice.

Benefits of technology

It improved the effectiveness of the prompting function in video recording, reduced the risk of forgetting lines, and enhanced the quality of recorded videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114911448B_ABST
    Figure CN114911448B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method, device and equipment and medium, the method comprises: in response to the service start operation in the video application, starting the video recording service in the video application; collecting user voice in the video recording service, determining the target text associated with the user voice in the prompt text data associated with the video recording service, highlighting the target text; when the text position of the target text in the prompt text data is the end position in the prompt text data, obtaining the target video data corresponding to the video recording service. By using the embodiments of the present application, the effectiveness of the word prompt function in the video recording service can be improved, and the quality of the recorded video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet, and particularly relates to a data processing method and device, equipment and medium. BACKGROUND

[0002] With the development of short videos, more and more users (including people without any shooting and editing experience) join the ranks of multimedia creators and start to show their performances in front of the camera. For multimedia creators lacking experience, forgetting lines and other situations often occur in front of the camera, and even if the content script is memorized, there will be problems such as stumbling or unnatural expressions.

[0003] In the prior art, in the process of shooting a short video, the user can print the content of the script and place it beside the camera for prompting; however, when the content of the script is too much, the user may not be able to quickly locate the content to be delivered, or there may be a positioning error, and the effect of delivering lines by printing the content of the script is not obvious, and when the user looks at the script content beside the camera, the action of the user will be shot by the camera, thereby affecting the quality of the final shot video. SUMMARY

[0004] The embodiments of the present application provide a data processing method, device, equipment and medium, which can improve the effectiveness of the teleprompter function in the video recording service, and thereby improve the quality of the recorded video.

[0005] The embodiments of the present application provide a data processing method, device, equipment and medium, which can improve the effectiveness of the teleprompter function in the video recording service, and thereby improve the quality of the recorded video.

[0006] In response to a service start operation in a video application, a video recording service in the video application is started;

[0007] User speech in the video recording service is collected, target text associated with the user speech is determined in prompt text data associated with the video recording service, and the target text is highlighted;

[0008] When the text position of the target text in the prompt text data is the end position of the prompt text data, target video data corresponding to the video recording service is obtained.

[0009] The embodiments of the present application provide a data processing method, device, equipment and medium, which can improve the effectiveness of the teleprompter function in the video recording service, and thereby improve the quality of the recorded video.

[0010] The prompt text data is uploaded to a teleprompter application;

[0011] User speech corresponding to the target user is collected, and the user speech is converted into text to generate user speech text corresponding to the user speech;

[0012] In the prompt text data, the same text as the user voice text is determined as the target text, and the target text is highlighted in the teleprompter application.

[0013] The embodiment of the application provides a data processing device, comprising:

[0014] The starting module is configured to start a video recording service in the video application in response to a service starting operation in the video application.

[0015] The display module is configured to collect a user voice in the video recording service, determine a target text associated with the user voice in prompt text data associated with the video recording service, and highlight the target text.

[0016] The acquisition module is configured to acquire target video data corresponding to the video recording service when a text position of the target text in the prompt text data is an end position in the prompt text data.

[0017] The device further comprises:

[0018] The first recording page display module is configured to display a recording page in the video application in response to a triggering operation on a teleprompter shooting entrance in the video application. The recording page comprises a text input area.

[0019] The editing module is configured to display prompt text data determined by an information editing operation on the text input area in the text input area in response to the information editing operation.

[0020] The first estimated duration display module is configured to display a number of prompt characters and a video estimated duration corresponding to the prompt text data in the text input area when a number of prompt characters corresponding to the prompt text data is greater than a number threshold.

[0021] The device further comprises:

[0022] The second recording page display module is configured to display a recording page in the video application in response to a triggering operation on a teleprompter shooting entrance in the video application. The recording page comprises a text upload control and a text input area.

[0023] The text upload module is configured to determine text content uploaded to the recording page as prompt text data in response to a triggering operation on the text upload control, and display the prompt text data in the text input area.

[0024] The second estimated duration display module is configured to display a number of prompt characters corresponding to the prompt text data and a video estimated duration corresponding to the prompt text data.

[0025] The service starting operation comprises a voice starting operation.

[0026] The starting module comprises:

[0027] The countdown animation display unit is configured to display a recording countdown animation associated with the video recording service in a recording page of the video application in response to the voice starting operation in the video application.

[0028] The recording service starting unit is configured to start and execute the video recording service in the video application when the recording countdown animation ends.

[0029] The recording countdown animation comprises an animation cancel control.

[0030] The device further comprises:

[0031] The countdown animation cancel module is configured to cancel the display of the recording countdown animation and start and execute the video recording service in the video application in response to a triggering operation on the animation cancel control.

[0032] The display module comprises:

[0033] The voice endpoint detection unit is configured to collect initial user voice in the video recording service, perform voice endpoint detection on the initial user voice, and determine valid voice data in the initial user voice as user voice.

[0034] The target text determination unit is configured to convert the user voice into user voice text, perform text matching on the user voice text and prompt text data associated with the video recording service, and determine target text in the prompt text data that matches the user voice text.

[0035] The target text display unit is configured to highlight the target text in a recording page of the video recording service.

[0036] The target text determination unit comprises:

[0037] The syllable information acquisition subunit is configured to acquire first syllable information corresponding to the user voice text and second syllable information corresponding to the prompt text data associated with the video recording service.

[0038] The syllable matching subunit is configured to acquire target syllable information identical to the first syllable information in the second syllable information and determine target text corresponding to the target syllable information in the prompt text data.

[0039] The target text display unit comprises:

[0040] The prompt area determination subunit is configured to determine a text prompt area corresponding to the target text in a recording page of the video recording service.

[0041] The highlighting subunit is configured to highlight the target text in the text prompt area according to the text position of the target text in the prompt text data.

[0042] The recording page includes a recording cancellation control;

[0043] The device further includes:

[0044] The recording cancellation module is configured to cancel the video recording service and delete video data recorded by the video recording service in response to a triggering operation on the recording cancellation control.

[0045] The recording prompt information display module is configured to generate recording prompt information for the video recording service and display the recording prompt information in the recording page. The recording prompt information includes a re-recording control.

[0046] The re-recording module is configured to switch the target text displayed in the recording page to the prompt text data in response to a triggering operation on the re-recording control.

[0047] The recording page includes a recording completion control;

[0048] The device further includes:

[0049] The recording completion module is configured to stop the video recording service and determine video data recorded by the video recording service as target video data in response to a triggering operation on the recording completion control.

[0050] The obtaining module includes:

[0051] The original video obtaining unit is configured to stop the video recording service and determine video data recorded by the video recording service as original video data when the text position of the target text in the prompt text data is an end position in the prompt text data.

[0052] The optimization control display unit is configured to display the original video data and a clip optimization control corresponding to the original video data in a clip page of the video application.

[0053] The optimization mode display unit is configured to display M clip optimization modes for the original video data in response to a triggering operation on the clip optimization control. M is a positive integer.

[0054] The optimization processing unit is configured to perform clip optimization processing on the original video data according to a clip optimization mode determined by a selection operation on the M clip optimization modes to obtain target video data corresponding to the video recording service.

[0055] The optimization processing unit includes:

[0056] The first speech conversion subunit is configured to, if the clip optimization mode determined by the selection operation is a first clip mode, obtain target speech data contained in the original video data, and convert the target speech data into a target text result.

[0057] The text comparison subunit is configured to compare the target text result with the prompt text data, and determine text in the target text result that is different from the prompt text data as error text.

[0058] The speech deletion subunit is configured to delete speech data corresponding to the error text in the original video data, and obtain target video data corresponding to the video recording service.

[0059] The optimization processing unit comprises:

[0060] The second speech conversion subunit is configured to, if the clip optimization mode determined by the selection operation is a second clip mode, convert target speech data contained in the original video data into a target text result, and determine text in the target text result that is different from the prompt text data as error text.

[0061] The timestamp acquisition subunit is configured to divide the target text result into N text characters, and acquire timestamps of the N text characters in the target speech data; N is a positive integer.

[0062] The speech pause segment determination subunit is configured to determine a speech pause segment in the target speech data according to the timestamps, delete speech data corresponding to the speech pause segment and the error text in the original video data, and obtain target video data corresponding to the video recording service.

[0063] The device further comprises:

[0064] The user speech speed determination module is configured to acquire a speech duration corresponding to the initial speech of the user, and a number of speech characters contained in the initial speech of the user, and determine a ratio of the number of speech characters to the speech duration as the user speech speed.

[0065] The speech speed prompt information display module is configured to display speech speed prompt information in the recording page when the user speech speed is greater than a speech speed threshold; the speech speed prompt information is used to prompt the target user associated with the video recording service to reduce the user speech speed.

[0066] The error text comprises K error subtexts, and K is a positive integer.

[0067] The device further comprises:

[0068] The error frequency determination module is configured to determine an error frequency in the video recording service according to the K error subtexts and a video duration corresponding to the original video data.

[0069] The error type identification module is configured to identify the speech error types corresponding to the K error subtexts respectively when the error frequency is greater than the error threshold.

[0070] The course video pushing module is configured to push the course video associated with the speech error type to the target user associated with the video recording service in the video application.

[0071] The application embodiment provides a data processing apparatus, including:

[0072] The prompt text uploading module is configured to upload the prompt text data to the teleprompter application.

[0073] The user voice collecting module is configured to collect the user voice corresponding to the target user, convert the user voice into text, and generate the user voice text corresponding to the user voice.

[0074] The user voice text display module is configured to determine the same text as the user voice text as the target text in the prompt text data, and highlight the target text in the teleprompter application.

[0075] The target user includes a first user and a second user, and the prompt text data includes a first prompt text corresponding to the first user and a second prompt text corresponding to the second user.

[0076] The user voice text display module includes:

[0077] The user identity determination unit is configured to obtain the user voiceprint feature in the user voice, and determine the user identity corresponding to the user voice according to the user voiceprint feature.

[0078] The first determination unit is configured to, if the user identity is the first user, determine the same text as the user voice text as the target text in the first prompt text, and highlight the target text in the teleprompter application.

[0079] The second determination unit is configured to, if the user identity is the second user, determine the same text as the user voice text as the target text in the second prompt text, and highlight the target text in the teleprompter application.

[0080] The application embodiment provides a computer device, including a memory and a processor, the memory is connected with the processor, the memory is used for storing a computer program, and the processor is used for calling the computer program to make the computer device execute the method provided in the above-mentioned aspect of the application embodiment.

[0081] The embodiment of the application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is suitable for being loaded and executed by a processor, so that a computer device with the processor executes the method provided in the above aspect of the embodiment of the application.

[0082] According to an aspect of the application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the method provided in the above aspect.

[0083] The embodiment of the application can start a video recording service in a video application in response to a service starting operation, collect user voice in the video recording service, determine target text associated with the user voice in prompt text data associated with the video recording service, highlight the target text, and obtain target video data corresponding to the video recording service when a text position of the target text in the prompt text data is an end position in the prompt text data. It can be seen that after starting the video recording service in the video application, the target text matched with the user voice can be located in the prompt text data, and the target text is highlighted in the video application. The target text displayed in the video application matches the content of the speech of the user, which can improve the effectiveness of the text prompt function of the video recording service, reduce the risk of recording failure caused by forgetting words of the user, and further improve the quality of the recorded video. BRIEF DESCRIPTION OF DRAWINGS

[0084] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed in the embodiment or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.

[0085] Figure 1 is a structural schematic diagram of a network architecture provided by the embodiment of the application;

[0086] Figure 2 is a schematic diagram of a data processing scene provided by the embodiment of the application;

[0087] Figure 3 is a flowchart of a data processing method provided by the embodiment of the application;

[0088] Figure 4 is an interface schematic diagram of input prompt text data provided by the embodiment of the application;

[0089] Figure 5 is a kind of interface schematic diagram provided by the video recording service start in the video application of an embodiment of the present application;

[0090] Figure 6 is a kind of interface schematic diagram provided by the display prompt text data of an embodiment of the present application;

[0091] Figure 7 is a kind of interface schematic diagram provided by the display speech speed prompt information of an embodiment of the present application;

[0092] Figure 8 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0093] Figure 9 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0094] Figure 10 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0095] Figure 11 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0096] Figure 12 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0097] Figure 13 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0098] Figure 14 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0099] Figure 15 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0100] Figure 16 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application;

[0101] Figure 17 is a kind of interface schematic diagram provided by the video recording service stop of an embodiment of the present application. DETAILED DESCRIPTION

[0102] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0103] Please refer to Figure 1 , Figure 1 is a structural schematic diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, the network architecture can include a server 10d and a user terminal cluster, which can include one or more user terminals, and the number of user terminals is not limited here. As Figure 1 shown, the user terminal cluster can specifically include a user terminal 10a, a user terminal 10b, and a user terminal 10c, etc. The server 10d can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform, etc. The user terminal 10a, the user terminal 10b, and the user terminal 10c, etc. can each include a smart phone, a tablet computer, a notebook computer, a palm computer, a mobile internet device (MID), a wearable device (such as a smart watch, a smart bracelet, etc.), and a smart television, etc. smart terminal with video / image playing function. As Figure 1 shown, the user terminal 10a, the user terminal 10b, and the user terminal 10c, etc. can be respectively connected with the server 10d in a network, so that each user terminal can interact with the server 10d through the network connection.

[0104] In Figure 1Taking the user terminal 10a as an example, the user terminal 10a can be equipped with a video application with video recording capabilities. This video application can be a video editing application, a short video application, etc. The user can open the video application installed on the user terminal 10a. This video application can provide the user with video recording capabilities, which can include conventional shooting methods and prompting shooting methods. Conventional shooting methods refer to the process of filming the user using the built-in camera of the user terminal 10a (or an external camera device that has a communication connection with the user terminal 10a). In this case, it may not be possible to provide the user with scripted content prompts, and the user needs to prepare the scripted content they want to express during the video recording process in advance (for example, write down the scripted content). Prompting shooting methods refer to the process of filming the user using the built-in camera of the user terminal 10a or an external camera device. In this case, the scripted content can be displayed on the terminal screen of the user terminal 10a, and the display of the scripted content can be switched according to the user's speech progress (for example, scrolling display). The scripted content here can also be referred to as the prompt text data in the video recording. After a user triggers the prompting entry point (i.e., the prompting entry point) in a video application, the user terminal 10a can respond to the triggering operation and display the recording page in the video application. Before recording, the user can input prompt text data on the recording page or upload existing prompt text data to the recording page. When the user starts video recording, the user terminal 10a can respond to the user's video recording start operation, start the video recording function in the video application, and display prompt text data on the user terminal 10a's screen according to the user's speech progress during video recording. In other words, during video recording, prompt text data can be displayed according to the user's speech progress. When the user's speech speed increases, the switching speed (which can be scrolling speed) of the prompt text data in the video application increases; when the user's speech speed decreases, the switching speed of the prompt text data in the video application decreases. That is, the text displayed in the video application matches the user's speech to ensure the effectiveness of the prompt content during video recording, helping the user to smoothly complete the video recording, thereby improving the quality of the recorded video.

[0105] Please see also Figure 2 , Figure 2 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Taking a video recording scenario as an example, the implementation process of the data processing method provided in this embodiment of the application is described. Figure 2 The user terminal 20a shown can be the one described above. Figure 1Any one of the user terminals in the illustrated user terminal cluster, the user terminal 20a is installed with a video application, and the video application has a video recording function. User A (the user A can refer to the user of the user terminal 20a) can open the video application in the user terminal 20a and enter the home page of the video application, the user can perform a trigger operation on the shooting entrance in the video application, and the user terminal 20a displays a shooting page 20m in the video application in response to the trigger operation on the shooting entrance. The shooting page 20m can include a shooting area 20b, a filter control 20c, a shooting control 20d, and a beautification control 20e, etc. Among them, the shooting area 20b is used to display the video picture collected by the user terminal 20a, which can be the video picture of the user A, and can be collected by the camera of the user terminal 20a or the camera device having communication connection with the user terminal 20a; the shooting control 20d can be used to control the start and stop of video recording, and performing a trigger operation on the shooting control 20d after entering the shooting page 20m can represent starting shooting, and the video picture shot can be displayed in the shooting area 20b. When performing a trigger operation on the shooting control 20d again during the shooting process, it can represent stopping shooting, and the video picture displayed in the shooting area 20b will be fixed at the picture when stopping shooting; the filter control 20c can be used to perform image processing on the video picture collected by the user terminal 20a to achieve a certain special effect, such as the skin filter which can perform skin modification, color polishing and skin polishing on the portrait in the collected video picture; the beautification control 20e can be used to perform beautification processing on the portrait in the video picture collected by the user terminal 20a, such as automatically repairing the face shape of the portrait, increasing the eyes of the portrait, and changing the nose of the portrait, etc.

[0106] In the shooting page 20m, a word prompting shooting entrance 20f can also be included. When the user A lacks experience in recording a video, in order to prevent the situation of forgetting words during video recording (if words are forgotten, the video may need to be re-recorded), the user A can select the word prompting shooting function in the video application, that is, the user A can perform a trigger operation on the word prompting shooting entrance 20f in the shooting page 20m, and the user terminal 20a can respond to the trigger operation of the user A on the word prompting shooting entrance 20f to switch the shooting page 20m in the video application to a recording page corresponding to the word prompting shooting entrance 20f. The recording page can first display a text input area, and the user A can input the script content required for recording the video in the text input area. The script content can be used to prompt the user A during video recording. In short, during the recording of the video, the user A can record according to the script content displayed in the video application. At this time, the script content can also be referred to as prompt text data 20g. In the text input area, statistical information 20h of the script content input by the user A can also be displayed. The statistical information 20h can include the number of words in the input script content (for example, the number of words of the script content is 134), and the video estimated time corresponding to the input script content (such as 35 seconds). The user A can increase or delete the script content according to the video estimated time. For example, the user A wants to record a 1-minute video. When the video estimated time corresponding to the script content input by the user A in the text input area is 4 minutes, the user A can delete the script content displayed in the text input area so that the video estimated time corresponding to the deleted script content is about 1 minute (for example, the video estimated time can range from 55 seconds to 65 seconds). When the video estimated time corresponding to the script content input by the user A in the text input area is 35 seconds, the user A can increase the script content displayed in the text input area so that the video estimated time corresponding to the increased script content is about 1 minute. Then, the user A can determine the final script content as the prompt text data 20g.

[0107] Further, after the user A determines the prompt text data 20g, the user A can perform a trigger operation on the "next step" control in the recording page, and the user terminal 20a can respond to the trigger operation on the "next step" control to start the camera (or the camera device with communication connection) of the user terminal 20a, and enter a video recording preparation state (i.e., before the video starts recording); for example, Figure 2As shown, the video picture 20i collected by the user terminal 20a for the user A can be displayed in the recording page, and the prompt information "adjust the position and place the mobile phone, say'start' to start the teleprompting shooting" can be displayed in the recording page, that is, the user A can adjust the position of the user A and the position of the user terminal 20a according to the video picture 20i, and after the position is adjusted, the user can start the video recording through voice, for example, the user can say "start" to start the video recording.

[0108] After the user A says "start", the user terminal 20a can start the video recording in the video application in response to the voice starting operation of the user A, and display the prompt text data 20g in the recording page. It can be understood that the text displayed in the recording page can be only part of the prompt text data 20g, for example, a sentence in the prompt text data 20g, so the first sentence in the prompt text data 20g can be displayed first after the video recording is started. When the user A starts the speech during the video recording, the user terminal 20a can collect the user voice corresponding to the user A, and the client of the video application installed in the user terminal 20a can transmit the user voice to the background server 20j of the video application and send a voice matching instruction to the background server 20j. After receiving the user voice and the voice matching instruction, the background server 20j can convert the user voice into user voice text, and when the user voice text is Chinese (at this time, it can be defaulted that the prompt text data 20g is also Chinese), the background server 20j can also convert the user voice text into first Chinese pinyin; of course, after the user A inputs the prompt text data 20g in the text input area, the client of the video application can also transmit the prompt text data 20g to the background server 20j, so the background server 20j can convert the prompt text data 20g into second Chinese pinyin. The background server 20j can match the first Chinese pinyin with the second Chinese pinyin, find the same pinyin in the second Chinese pinyin as the first Chinese pinyin, that is, find the text position of the first Chinese pinyin in the second Chinese pinyin, determine the text corresponding to the text position in the prompt text data 20g as the target text (that is, the text matched by the user voice in the prompt text data 20g), and transmit the target text to the client of the video application, so that the terminal device 20a can highlight the target text in the video application (for example, increase the display size of the target text, change the display color of the target text, etc.). It can be understood that when the user A speaks according to the order of the text prompt data, the prompt text data can be displayed in the recording page; when the user A does not speak according to the order of the text prompt data, the prompt text data can be displayed in the recording page.

[0109] Optionally, when the target text is a word or a phrase, the sentence where the target text is located can be highlighted in the video application. For example,Figure 2 As shown, when the user voice is "weekend", the background server 20j can match the target text corresponding to the user voice in the prompt text data 20g as "weekend", and at this time, the sentence "weekend, attend the consumption class in Changsha for cooperation with xx and xx" in which the target text "weekend" is located can be highlighted (increase the text display size and bold the text, as shown in the area 20k in Figure 2 ).

[0110] It should be noted that the prompt text data 20g can be directly displayed in the recording page, or can be displayed in a sub-page displayed independently of the recording page. The application does not limit the display form of the prompt text data 20g in the recording page. The purpose of matching the user voice in the prompt text data 20g is to determine the text position of the user voice in the prompt text data 20g. When converting the user voice into user voice text, only the consistency between the character pronunciation and the user voice needs to be considered, and the accuracy between the converted user voice text and the user voice does not need to be considered. Therefore, the Chinese audio can be matched, and the matching efficiency between the user voice and the prompt text data can be improved.

[0111] The user terminal 20a can collect the user voice of user A in real time, the background server 20n can determine the target text corresponding to the user voice in the prompt text data 20g in real time, and then the text prompt information can be displayed in real time according to the user voice progress. For example, when user A speaks the first sentence in the prompt text data 20g, the first sentence in the prompt text data 20g can be highlighted in the recording page; when user A speaks the second sentence in the prompt text data 20g, the recording page can be switched from the first sentence to the second sentence in the prompt text data 20g, and the second sentence can be highlighted. The target text highlighted in the recording page is the content of the current speech of user A. When user A speaks the last word in the prompt text data 20g, the user terminal 20a can close the video recording, and determine the recorded video as a recording completed video. If user A is satisfied with the recorded video, the video can be saved; if user A is not satisfied with the recorded video, user A can re-shoot. Of course, user A can also edit and optimize the recorded video to obtain a final recording video.

[0112] As shown in the video recording process of the application embodiment, the prompt text data can be displayed according to the user voice progress to achieve the precise word prompt effect of the user, and thus the quality of the recorded video can be improved.

[0113] Please refer to Figure 3 , Figure 3is a flowchart of a data processing method provided by an embodiment of the present application. It can be understood that the data processing method can be executed by a computer device, which can be a user terminal, or an independent server, or a cluster composed of multiple servers, or a system composed of a user terminal and a server, or a computer program application (including program code), which is not limited here. As shown in Figure 3 The data processing method can include the following steps S101-S103:

[0114] Step S101, in response to a service start operation in a video application, starting a video recording service in the video application.

[0115] Specifically, when a user wants to express his own views or show his own life in front of the camera, he can record a video in the video application to record the video he wants. For the final recorded video, the user can upload it to an information publishing platform for sharing, so that users in the information publishing platform can watch the recorded video. In the embodiment of the present application, the user who needs to record a video can be referred to as a target user, and the device used by the target user when recording a video can be referred to as a computer device. When the target user executes a service start operation for the video recording service in the video application installed on the computer device, the computer device can start the video recording service in the video application in response to the service start operation in the video application, that is, start video recording in the video application. The service start operation can include but is not limited to touch trigger operations such as single click, double click, long press, screen tapping, and non-contact trigger operations such as voice, remote control, and gestures.

[0116] Among them, before the computer device starts the video recording service, the target user also needs to upload the prompt text data required in the video recording service to the video application. The prompt text data can be used to prompt the target user in the video recording service, which can greatly reduce the situation of forgetting words in the video recording process of the target user. After the target user opens the video application installed in the computer device, he can enter the shooting page in the video application (for example, the shooting page in the video application shown in FIG. 1A). Figure 2In the shooting page 20m) in the corresponding embodiment, the shooting page of the video application can include a teleprompter shooting entry. When the target user performs a trigger operation on the teleprompter shooting entry in the shooting page, the computer device can display a recording page in the video application in response to the trigger operation on the teleprompter shooting entry in the video application, and the recording page can include a text input area that can be used to edit text content. The computer device can display the prompt text data determined by the information editing operation in the text input area in response to the information editing operation on the text input area. When the number of prompt words corresponding to the prompt text data is greater than a number threshold (the number threshold can be pre-set according to actual needs, for example, the number threshold can be set to 100), the video estimated duration corresponding to the number of prompt words and the number of prompt text can be displayed in the text input area. In other words, after the target user performs a trigger operation on the teleprompter shooting entry in the shooting page, the shooting page can be switched to display the recording page in the video application, and the target user can edit the script content (i.e., the above-mentioned prompt text data) needed in the video recording service in the text input area of the recording page. When the target user edits the text in the text input area, the number of prompt words input in the text input area can be counted and displayed in real time. When the number of prompt words is greater than the pre-set number threshold, the video estimated duration corresponding to the currently input prompt text data can be displayed in the text input area. Optionally, the teleprompter shooting entry can be displayed in any page of the video application in addition to the shooting page. The display position of the teleprompter shooting entry is not limited in the embodiments of the present application.

[0117] The video estimated duration can be used as the duration reference information of the finished video recorded in the subsequent video recording service. When there is a large difference between the video estimated duration displayed in the text input area and the video duration expected by the target user, the target user can increase or delete the text in the text input area. For example, when the video estimated duration displayed in the text input area is 35 seconds and the target user expects to record a video of 2 minutes, the target user can continue to edit the text in the text input area until the video estimated duration displayed in the text input area is within a set duration range (for example, the video estimated duration is between 1 minute 50 seconds and 2 minutes 10 seconds).

[0118] Optionally, the computer device can also display a text uploading control on the recording page in response to the triggering operation on the word suggestion shooting entry in the video application, and the target user can perform a triggering operation on the text uploading control on the recording page to upload the edited prompt text data to the recording page, i.e., the computer device can determine the text content uploaded to the recording page as the prompt text data in response to the triggering operation on the text uploading control, display the prompt text data in the text input area of the recording page, and also display the number of prompt words corresponding to the prompt text data and the video estimated duration corresponding to the prompt text data. The text uploading control can include but is not limited to a text pasting control and a last text selection control. When the target user performs a triggering operation on the text pasting control, it means that the target user can directly paste the pre-edited prompt text data into the text input area without temporarily editing the text content. When the target user performs a triggering operation on the last text selection control, it means that the target user can use the prompt text data in the last video recording service in the current video recording service, i.e., the target user may not be satisfied with the finished video recorded in the last video recording service, so the target user can re-record in the current video recording service to avoid repeated input of the same prompt text data, thereby improving the input efficiency of the prompt text data.

[0119] Please refer to Figure 4 , Figure 4 is an interface schematic diagram provided by an embodiment of the present application for inputting prompt text data. As shown in Figure 4 , after the target user performs a triggering operation on the shooting entry in the video application installed on the user terminal 30a, the user terminal 30a can display a shooting page 30g in the video application in response to the triggering operation on the shooting entry (the user terminal 30a at this time can be the computer device described above), which can include a shooting area 30b, a filter control 30c, a shooting control 30d, a beautification control 30e, and a word suggestion shooting entry 30f, etc. The functions of the shooting area 30b, the filter control 30c, the shooting control 30d, and the beautification control 30e in the video application are described in the above Figure 2 corresponding embodiments, and will not be described here.

[0120] When the target user performs a triggering operation on the text shooting entry 30f in the shooting page 30g, the user terminal 30a can respond to the triggering operation on the text shooting entry 30f in the shooting page 30g, and switch the shooting page 30g to a recording page 30h in the video application. The recording page 30h can include a text input area 30i, which can be used to directly edit the text content. The target user can click the text input area 30i to pop up a keyboard 30p in the recording page 30h. The target user can edit the prompt text data needed in the current video recording service through the keyboard 30p. The user terminal 30a can respond to the information editing operation of the target user, and display the text content determined by the information editing operation in the text input area 30i. At the same time, the user terminal 30a can count the number of characters of the text content input in the text input area 30i in real time. When the number of characters of the text content input in the text input area 30i is greater than a pre-set number threshold (for example, the number threshold is set to 100), the user terminal 30a can display the number of characters of the input text content and the estimated shooting time (i.e., the video estimated time) corresponding to the input text content in the area 30m of the text input area 30i. For example, as shown in Figure 4 When the target user inputs the text content "weekend, participate in xx and xx's consumption class in Changsha. When the year, others are through the public number online" in the text input area 30i, the user terminal 30a counts the number of characters of the text content as 32, and the estimated shooting time is 15 seconds, that is, "current number of characters 32, estimated shooting time 15 seconds" is displayed in the area 30m. The target user can edit the text content according to the estimated shooting time displayed in the area 30m. After the target user completes the editing of the text content in the text input area 30i, the target user can determine the text content in the text input area 30i as the prompt text data, and then perform a triggering operation on the "next step" control 30n in the recording page 30h to trigger the user terminal 30n to enter the next step operation of the video recording service.

[0121] Optionally, as Figure 4As shown, the text input area 30i can further include a paste text control 30j and a last text control 30k. When the target user performs a triggering operation on the paste text control 30j, it indicates that the target user has edited the prompt text data in the remaining application and copied the prompt text data from the remaining application. In response to the triggering operation on the paste text control 30j, the user terminal 30a pastes the prompt text data copied by the target user into the text input area 30i. When the video recorded by the target user in the current video recording service is a re-recording of the video recorded in the last video recording service, the target user can perform a triggering operation on the last text control 30k. In response to the triggering operation on the last text control 30k, the user terminal 30a acquires the prompt text data in the last video recording service and displays the prompt text data in the text input area 30i, directly using the prompt text data used in the last video recording service as the prompt text data in the current video recording service. Optionally, the target user can adjust the prompt text data used in the last video recording service in the text input area 30i according to the experience of the last video recording service. For example, the target user finds that the statement 1 in the prompt text data has a logical error in the last video recording service, and in the current video recording service, the target user can modify the prompt text data of the last video recording service in the text input area 30i.

[0122] It should be noted that the number of characters and the pre-fragmentation duration of the prompt text data input into the text input area 30i through the paste text control 30j and the last text control 30k can also be displayed in the area 30m of the text input area 30i. In the embodiment of the present application, the target user inputs the prompt text data in the video recording service into the text input area 30i by using the paste text control 30j and the last text control 30k, which can improve the input efficiency of the prompt text data in the video recording service.

[0123] Optionally, when the service starting operation is a voice starting operation, the target user can perform a voice starting operation on the video recording service in the video application after completing the editing operation of the prompt text data, and the computer device can display a recording countdown animation associated with the video recording service in the recording page of the video application in response to the voice starting operation. When the recording countdown animation ends, the video recording service in the video application is started and executed, that is, the video recording is formally started. During the playing of the recording countdown animation in the recording page, the camera device corresponding to the computer device can be started, and the target user can adjust the position of himself and the computer device according to the video image displayed in the recording page to find the best shooting angle. Optionally, the recording countdown animation corresponding animation cancel control can also be displayed in the recording page. When the target user has prepared for recording the video, a trigger operation can be performed on the animation cancel control to cancel the recording countdown animation; that is, the computer device can cancel the display of the recording countdown animation in the recording page in response to the trigger operation of the target user on the animation cancel control, and start and execute the video recording service in the video application. In other words, after the target user starts the video recording service by voice, the target user will not directly enter the formal recording mode in the video application, but the recording countdown animation will be played in the recording page to provide the target user with a short recording preparation time (that is, the length of the recording countdown animation, such as 5 seconds), and the target user will enter the formal recording mode after the playing of the recording countdown animation is completed; or the target user can also cancel the playing of the recording countdown animation and directly enter the formal recording mode under the premise of preparing for recording in advance.

[0124] Please see Figure 5 , Figure 5 is an interface schematic diagram provided by an embodiment of the present application for starting a video recording service in a video application. After completing the editing operation of the prompt text data, the target user can perform the next operation (such as performing a trigger operation on the “next step” control 30n in the embodiment corresponding to the above Figure 4 . Figure 5As shown, after the target user edits the prompt text data and performs the next step operation, the target user can exit the text input area in the recording page 40b, display the video picture of the target user in the area 40c of the recording page 40b, and at the same time, display the prompt information 40d (adjust the position, place the mobile phone, say "start" to start the teleprompter shooting) in the recording page 40b, that is, before starting the video recording service, the user terminal 40a (the user terminal 40a at this time can be referred to as a computer device) can start the camera device (such as the camera of the user terminal 40a) associated with it, collect image data of the target user, and render the collected image data into a video picture corresponding to the target user, and display the video picture of the target user in the area 40c of the recording page 40b. The target user can adjust the position of himself and the lens according to the video picture displayed in the area 40c to find the best shooting angle.

[0125] After the target user adjusts the position of himself and the lens, that is, after the target user completes the preparation work for recording the video, the target user can say "start" to start the video recording service in the video application. When the target user says "start" to perform the voice start operation on the video recording service in the video application, the user terminal 40a can display a recording countdown animation in the area 40e of the recording page 40b in response to the voice start operation on the video recording service. The length of the recording countdown animation can be 5 seconds. Of course, the first few sentences (such as the first two sentences) of the prompt text data can also be displayed in the area 40e of the recording page 40b.

[0126] Further, when the recording countdown animation in the recording page 40b ends, the user terminal 40a can start and execute the video recording service in the video application. Alternatively, if the target user does not want to wait for the recording countdown animation to end before starting the video recording service, the target user can perform a trigger operation on the animation cancel control 40f in the recording page 40b to cancel the recording countdown animation in the recording page 40b and directly start and execute the video recording service. After starting the formal video recording, the target user can start speaking, and the user terminal 40a can collect the user voice of the target user, find the target text matching the user voice in the prompt text data, and highlight the target text (such as bolding and enlarging the target text) in the area 40g of the recording page 40b. The specific determination process of the target text will be described in the following step S102.

[0127] Step S102: Collecting the user voice in the video recording service, determining the target text associated with the user voice in the prompt text data associated with the video recording service, and highlighting the target text.

[0128] Specifically, after starting formal video recording, the computer device can start the audio collection function, collect the user voice of the target user in the video recording service, and find the target text matching the user voice in the prompt text data, and highlight the target text contained in the prompt text data in the recording page. The computer device can collect the user voice of the target user in the video recording service in real time, convert the user voice to text, determine the text position corresponding to the user voice in the prompt text data, determine the target text corresponding to the user voice according to the text position, and highlight the target text in the recording page. The highlighting can include but is not limited to text display color, text font size, and text background. The target text can refer to text data containing user voice converted text. For example, the user voice converted text is: New Year, and the target text can refer to a complete sentence containing "New Year", such as the target text is: At the arrival of the New Year, I wish you a prosperous New Year.

[0129] Further, the computer device directly collects the voice as the user initial voice, that is, the computer device can collect the user initial voice in the video recording service, perform voice activity detection (Voice Activity Detection, VAD) on the user initial voice, determine the valid voice data in the user initial voice as the user voice, and then convert the user voice to user voice text, perform text matching on the prompt text data associated with the video recording service, determine the target text matching the user voice text in the prompt text data, and highlight the target text in the recording page of the video recording service. In other words, the user initial voice collected by the computer device can contain noise of the environment where the target user is located and pause part in the speech process of the target user, so the user initial voice can be subjected to voice activity detection, the silence and noise in the user initial voice are deleted as interference information, and the valid voice data in the user initial voice is retained. The valid voice data can be referred to as the user voice of the target user. The computer device can convert the user voice to user voice text through a fast speech-to-text model, compare the user voice text with the prompt text data, find the text position of the user voice text in the prompt text data, and then determine the target text corresponding to the user voice in the text data according to the text position, and highlight the target text in the recording page of the video recording service.

[0130] The fast speech-to-text model refers to a process of converting user speech into text without context correction and without considering whether the semantics is correct, but only judging whether the pronunciation of the converted text is consistent with the user speech. When determining the target user matched with the user speech in the prompt text data, the computer device can determine the target text corresponding to the user speech in the prompt text data according to the pronunciation of the user speech text and the pronunciation of the prompt text data, that is, the computer device can obtain the first syllable information corresponding to the user speech text, obtain the second syllable information corresponding to the prompt text data associated with the video recording service, obtain the target syllable information same as the first syllable information in the second syllable information, and determine the target text corresponding to the target syllable information in the prompt text data.

[0131] The syllable information can refer to pinyin information in Chinese or phonetic symbol information in English, etc. When the prompt text data is Chinese, the computer device can convert the user speech text into first pinyin information, convert the prompt text data into second pinyin information, find the text position corresponding to the first pinyin information in the second pinyin information, and determine the target text corresponding to the user speech in the prompt text data according to the text position. When the prompt text data is English or other languages, the computer device can convert the user speech text into first phonetic symbol information, convert the prompt text data into second phonetic symbol information, and then determine the target text corresponding to the user speech in the prompt text data according to the first phonetic symbol information and the second phonetic symbol information. It can be understood that for Chinese, the same pronunciation can correspond to different characters, so the pinyin matching method can improve the efficiency of determining the target text. For languages in which different pronunciations correspond to different characters (for example, English), the computer device can directly match the letters contained in the user speech text with the letters contained in the prompt text data to determine the target text corresponding to the user speech in the prompt text data.

[0132] It should be noted that in the video recording service, the area for displaying the target text in the recording page can be set according to the terminal screen size of the computer device, such as the above Figure 5As shown in the area 40g in the recording page 40b, the display width of the area 40g is the same as the screen width of the computer device (such as the user terminal 40a), and the display height of the area 40g is less than the screen height of the computer device. When the terminal screen size of the computer device is large (such as the display screen of a desktop computer), if the size of the area for displaying the target text is the same as the terminal screen size of the computer device, the action of the target user in watching the target text (for example, the target user may move from the left to the right of the terminal screen when watching the target text) will be recorded in the video recording service, which causes the action and expression of the target user in the final recorded video to be unnatural, and further causes the quality of the recorded video to be too low. Therefore, in order to ensure that the action of the target user in the recorded video is natural, the text prompt area corresponding to the target text can be determined in the recording page of the video recording service according to the position of the camera corresponding to the computer device, and the target text can be highlighted in the text prompt area according to the text position of the target text in the prompt text data. In other words, in the video recording service, the target user can face the camera directly, and when the text prompt area and the camera of the computer device are located at the same position, the action of the target user in the video recorded by the video recording service is natural.

[0133] Please refer to Figure 6 , Figure 6 is a schematic diagram of an interface for displaying prompt text data provided by an embodiment of the present application. As shown in Figure 6 , after the user terminal 50a (i.e., the above-mentioned computer device) determines the target text "weekend, attend the consumption class in Changsha for cooperation between xx and xx" corresponding to the user voice in the prompt text data, the text prompt area 50e for displaying the target text can be determined in the recording page 50b of the video recording service according to the position of the camera 50d of the terminal device 50a, and the text prompt area 50e and the camera 50d are located at the same position. After starting to record the video formally, the video picture of the target user can be displayed in the area 50c of the recording page 50b, and the video recording time length (for example, the video recording time length is 00:13 seconds) can be displayed in the area 50f of the recording page 50b.

[0134] Optionally, in the video recording service, the computer device can collect the initial speech of the target user in real time, obtain the speech duration corresponding to the initial speech and the number of speech texts contained in the initial speech, determine the user speech speed as the ratio of the number of speech texts to the speech duration, and display the speech speed prompt information in the recording page when the user speech speed is greater than a speech speed threshold (the speech speed threshold can be artificially set based on actual needs, such as a speech speed threshold of 500 words per minute). The speech speed prompt information can be used to prompt the target user associated with the video recording service to reduce the user speech speed. In other words, the computer device can obtain the user speech speed of the target user in real time, and when the user speech speed is greater than the speech speed threshold, it indicates that the speech speed of the target user in the video recording service is too fast, and the target user can be reminded to appropriately slow down the speech speed.

[0135] Please refer to Figure 7 , Figure 7 is an interface schematic diagram provided by an embodiment of the present application for displaying speech speed prompt information. As shown in Figure 7 , after the user terminal 60a (i.e., the computer device) collects the initial speech of the target user, the user speech speed of the target user can be determined according to the number of speech texts contained in the initial speech and the speech duration. When the user speech speed of the target user in the video recording service is too fast (i.e., greater than the speech speed threshold), the speech speed prompt information 60c (for example, the speech speed prompt information can be “Your current speech speed is too fast. In order to ensure the quality of the recorded video, please slow down your speech speed”) can be displayed in the recording page 60b of the video recording service. Of course, in actual applications, the target user can also be reminded to slow down the speech speed in the form of voice broadcast, and the present application does not limit the display form of the speech speed prompt information.

[0136] Optionally, during the video recording process, the recording page of the video recording service can further include a cancel recording control and a complete recording control. After the target user performs a triggering operation on the cancel recording control in the recording page, the computer device can cancel the video recording service, delete the video data recorded by the video recording service, and generate recording prompt information for the video recording service in response to the triggering operation on the cancel recording control, and display the recording prompt information in the recording page. The recording prompt information can include a re-recording control. After the target user performs a triggering operation on the re-recording control, the computer device can switch the target text displayed by the recording page to the prompt text data in response to the triggering operation on the re-recording control, that is, display the text prompt information in the text input area of the recording page, and restart the video recording service. Of course, the recording prompt information can also include a return home control. After the target user performs a triggering operation on the return home control, the computer device can switch the recording page to the application home page in the video application in response to the triggering operation on the return home control, that is, cancel the video recording service being executed and temporarily stop starting the video recording service.

[0137] Optionally, after the target user performs a triggering operation on the complete recording control in the recording page, the computer device can stop the video recording service in response to the triggering operation on the complete recording control, and determine the video data recorded by the video recording service as the target video data recorded, that is, stop the video recording service before the prompt text data is completely presented, and the video recorded before the video recording service is stopped is referred to as the target video data.

[0138] Please see Figure 8 , Figure 8 is an interface schematic diagram for stopping a video recording service provided by an embodiment of the present application. As shown in Figure 8As shown, the user terminal 70a (i.e. the computer device described above) can determine the target text of the user voice in the prompt text data of the video recording service according to the user voice of the target user in the video recording service, highlight the target text in the recording page 70b, i.e. the user terminal 70a can display the prompt text data according to the progress of the user voice. During the video recording process, the recording page 70b can also display the cancel recording control 70c and the complete recording control 70d. When the target user performs a triggering operation on the complete recording control 70d, the user terminal 70a can stop the video recording service in response to the triggering operation on the complete recording control 70d, save the video data recorded in the video recording service, i.e. complete the video recording service; when the target user performs a triggering operation on the cancel recording control 70c, the user terminal 70a can cancel the video recording service in response to the triggering operation on the cancel recording control 70c, and delete the video data recorded in the video recording service. The user terminal 70a can generate recording prompt information 70e (for example, the recording prompt information can be "the recorded clip will be emptied, do you want to re-shoot the clip?") for the target user in the video recording service, and display the recording prompt information 70e in the recording page 70b of the video recording service. The recording prompt information 70e can include a "back to home page" control and a "re-shoot" control; when the target user performs a triggering operation on the "back to home page" control, the user terminal 70a can exit the video recording service and return to the application home page of the video application from the recording page 70b, i.e. the target user gives up re-shooting; when the target user performs a triggering operation on the "re-shoot" control, the user terminal 70a can exit the video recording service and return to the text input area from the recording page 70b, and display the prompt text data in the text input area, i.e. the target user chooses to re-record the video.

[0139] In step S103, when the text position of the target text in the prompt text data is the end position in the prompt text data, the target video data corresponding to the video recording service is obtained.

[0140] Specifically, in the video recording service, when the text position of the target text in the prompt text data is the end position in the prompt text data, it indicates that the target user has completed the shooting work of the video recording service, and the computer device can automatically end the video recording service and save the video data recorded in the video recording service without the target user's operation, and the video data recorded in the video recording service is determined as the target video data.

[0141] Optionally, the computer device can determine the video data saved when the video recording service is stopped as original video data, and enter a clip page of the video application. The original video data and clip optimization controls corresponding to the original video data are displayed in the clip page of the video application. The target user can perform a trigger operation on the clip optimization controls displayed in the clip page. In response to the trigger operation on the clip optimization controls, the computer device displays M clip optimization modes for the original video data, where M is a positive integer, that is, M can be 1, 2, …, and in this embodiment, the M clip optimization modes can include, but are not limited to, a clip optimization mode for removing errors (which can be referred to as a first clip mode) and a clip optimization mode for removing errors and sentence pauses (which can be referred to as a second clip mode). When the target user selects a clip optimization mode from the M clip optimization modes, the computer device can respond to the selection operation on the M clip optimization modes, and perform clip optimization processing on the original video data according to the clip optimization mode determined by the selection operation, to obtain target video data corresponding to the video recording service. It can be understood that the display area and display size of the original video data and the target video data in the clip page can be adjusted according to actual needs. For example, the display area of the original video data (or the target video data) can be located at the top of the clip page, or can be located at the bottom of the clip page, or can be located in the middle area of the clip page, etc. The display size of the original video data (or the target video data) can be a display ratio of 16:9, etc.

[0142] In the method, if the selected operation determines the first clipping mode, i.e., the target user selects the clipping optimization mode of removing errors, the computer device can obtain target voice data contained in the original video data, convert the target voice data into a target text result, and then perform text comparison between the target text result and the prompt text data, and determine the text in the target text result that is different from the prompt text data as error text. The computer device deletes the voice data corresponding to the error text in the original video data to obtain target video data corresponding to the video recording service. In the process of clipping optimization of the original video data, the computer device can use an accurate speech-to-text model to convert the target voice data contained in the original video data into text. The accurate speech-to-text model can learn semantic information in the target voice data, and needs to consider not only the consistency between the converted text and the user voice, but also the semantic information between the user voices, and correct the converted text through context semantic information. The computer device can perform voice endpoint detection on the target voice data contained in the original video data, remove noise and silence in the original video data, obtain effective voice data in the original video data, convert the effective voice data into text through the accurate speech-to-text model to obtain a target text result corresponding to the target voice data, and compare the text contained in the target text result with the text contained in the prompt text data one by one, and then determine the text different between the target text result and the prompt text data as error text. The error text can be caused by errors of the target user in the recording process of the video recording service. The computer device deletes the voice data corresponding to the error text from the original video data to obtain the final target video data.

[0143] Optionally, if the selected operation determines the second clipping mode, i.e., the target user selects the clipping optimization mode of removing errors and pauses between sentences, the computer device can convert target voice data contained in the original video data into a target text result, and determine the text in the target text result that is different from the prompt text data as error text. Then, the computer device can divide the target text result into N text characters, obtain time stamps of the N text characters in the target voice data, where N is a positive integer, such as 1, 2, …, and determine a voice pause segment in the target voice data according to the time stamps, delete the voice pause segment and the voice data corresponding to the error text in the original video data, and obtain target video data corresponding to the video recording service. The process of determining the error text by the computer device can refer to the description of the first clipping mode, which will not be repeated here.

[0144] The process of obtaining the speech pause segment by the computer device can include: the computer device can perform word segmentation processing on the target text result corresponding to the target speech data to obtain N text characters, obtain a timestamp of each text character in the target speech data, i.e., a timestamp in the original video data, obtain a time interval between each adjacent two text characters according to the timestamps corresponding to each adjacent two text characters in the N text characters, and determine a speech segment between the adjacent two text characters as a speech pause segment when the time interval between the adjacent two text characters is greater than a time threshold (for example, the time threshold can be set to 1.5 seconds). The number of speech pause segments can be one, multiple, or zero (i.e., no speech pause segment). For example, the N text characters can be represented in the order of arrangement in the target text result as: text character 1, text character 2, text character 3, text character 4, text character 5, and text character 6. The timestamp of text character 1 in the original video data is t1, the timestamp of text character 2 in the original video data is t2, the timestamp of text character 3 in the original video data is t3, the timestamp of text character 4 in the original video data is t4, the timestamp of text character 6 in the original video data is t5, and the timestamp of text character 6 in the original video data is t6. If the computer device calculates that the time interval between text character 2 and text character 3 is greater than the time threshold, the speech segment between text character 2 and text character 3 can be determined as speech pause segment 1. If it is calculated that the time interval between text character 5 and text character 6 is greater than the time threshold, the speech segment between text character 5 and text character 6 can be determined as speech pause segment 2. The speech corresponding to the error text and the video segments corresponding to speech pause segment 1 and speech pause segment 2 in the original video data are deleted to obtain the final target video data.

[0145] See Figure 9 , Figure 9 is an interface schematic diagram for optimizing video editing provided by an embodiment of the present application. As Figure 9As shown, after the video recording service is completed, the user can enter the editing page 80b of the video application, in which the video data 80c (e.g., the original video data as described above) recorded in the video recording service can be played back and previewed. The video data 80c can be displayed in the editing page 80b in a 16:9 ratio. A time axis 80d corresponding to the video data 80c can be displayed in the editing page 80b. The time axis 80d can include video nodes in the video data 80c, and the target user can quickly locate a play point in the video data 80c through the video nodes in the time axis 80d. The editing page 80b can further display an editing optimization control 80e (which can also be referred to as an editing optimization option button). When the target user performs a triggering operation on the editing optimization control 80e, the user terminal 80a (i.e., the computer device) can respond to the triggering operation on the editing optimization control 80e to pop up a selection page 80f in the editing page 80b (in an embodiment of the present application, the selection page can refer to a certain area in the editing page, or a sub-page displayed independently from the editing page, or a floating page in the editing page, or a page covering the editing page, and the display form of the selection page is not limited herein).

[0146] In the selection page 80f, different editing optimization modes for the video data 80c can be displayed, and the video time length corresponding to each editing optimization mode can be displayed. Figure 9As shown, if the target user selects "remove the part of the slip of the tongue" (i.e., the first clipping mode described above) in the selection page 80f, the video duration of the video data 80c after clipping optimization is 57 seconds (the video duration of the video data 80c is 60 seconds); if the target user selects "remove the slip of the tongue and the pause between sentences" (i.e., the second clipping mode described above) in the selection page 80f, the video duration of the video data 80c after clipping optimization is 50 seconds; if the target user selects no processing in the selection page 80f, i.e., the video data 80c is kept without processing. When the target user selects the optimization clipping mode of "remove the part of the slip of the tongue", the user terminal 80a can perform text conversion processing on the target voice data in the video data 80c to obtain the target text result corresponding to the target voice data, perform word matching between the target text result and the prompt text data, determine the error text, delete the voice data corresponding to the error text in the video data 80c, and obtain the target video data, where the target video data refers to the video data after deleting the part of the slip of the tongue. When the target user selects the optimization clipping mode of "remove the slip of the tongue and the pause between sentences", the user terminal 80a deletes the voice data corresponding to the error text in the video data 80c and the voice pause segment in the video data 80c, and further obtains the target video data, where the target video data refers to the video data after deleting the part of the slip of the tongue and the pause between sentences. After obtaining the target video data, the target user can save the target video data or upload the target video data to the information publishing platform, so that the user terminals in the information publishing platform can all watch the target video data.

[0147] Optionally, the error text can include K error subtexts, where K is a positive integer, such as 1, 2, …; the computer device can determine the error frequency in the video recording service according to the K error subtexts and the video duration corresponding to the original video data; when the error frequency is greater than an error threshold (for example, the error threshold can be set to 2 errors per minute), the computer device can identify the speech error types corresponding to the K error subtexts, and then can push the tutorial video associated with the speech error type to the target user associated with the video recording service in the video application. In other words, the computer device can recommend the corresponding tutorial video to the target user in the video application according to the speech error type corresponding to the error text, wherein the speech error type includes but is not limited to: non-standard Mandarin, pronunciation error, unclear pronunciation. For example, when the video duration of the original video data is 1 minute and the target user makes 3 errors in the original video data, the computer device can determine the speech error type of the error subtext corresponding to the 3 errors, and if the speech error type is a non-standard Mandarin type, the computer device can push a Mandarin tutorial video to the target user in the video application; if the speech error type is a pronunciation error type, the computer device can push a language tutorial video to the target user in the video application; if the speech error type is a unclear pronunciation type, the computer device can push a dubbing tutorial video to the target user in the video application.

[0148] See Figure 10 , Figure 10 is an interface schematic diagram for recommending a tutorial video according to a speech error type provided by an embodiment of the present application. As shown in Figure 10 , it is assumed that the target user selects the "remove the part of the slip of the tongue" clipping optimization mode to clip and optimize the original video data recorded in the video recording service, and obtains the target video data 90c after clipping optimization (i.e., the recorded video after removing the slip of the tongue); the user terminal 90a (i.e., the above-mentioned computer device) can display the target video data 90c in the clipping page 90b, and can also display a time axis 90d in the clipping page 90b, which can include video nodes associated with the target video data 90c. By triggering the video nodes in the time axis 90d, the target video data 90c at a specific time point can be located and played, and the target user can preview and play the target video data 90c in the clipping page 90b. The user terminal 90a can push the tutorial video matching the speech error type to the target user in the video application according to the speech error type corresponding to the error text in the clipping optimization process, as shown in Figure 10 , the speech error type corresponding to the error text is a non-standard Mandarin type, the user terminal 90a can obtain a tutorial video related to Mandarin teaching in the video application, and display the pushed Mandarin tutorial video in the area 90e of the editing page 90b.

[0149] Please refer to Figure 11 , Figure 11 is an implementation flowchart of a video recording service provided by an embodiment of the present application. As shown in Figure 11 , the implementation process of the video recording service is described by taking a client of a video application and a background server as examples. At this time, the client and the background server can be referred to as computer devices. The implementation flowchart of the video recording service can be implemented by the following steps S11-S25.

[0150] Step S11, input prompt text data, that is, the target user can open the client of the video application, enter the shooting page of the client, and enter the recording page from the word shooting entrance of the shooting page. At this time, the recording page includes a text input area, and the target user can input prompt text data in the text input area. After the target user completes the editing of the prompt text data, the target user can execute step S12, voice start "start", that is, "start" can be used as a wake-up word. After the target user says "start", the client can respond to the voice start operation of the user and execute step S13, start the video recording service, that is, start the recording mode.

[0151] Step S14, after entering the recording mode, the target user can read the text on the screen of the terminal device on which the client is installed (at this time, the text on the screen of the terminal device can be part of the text content in the prompt text data, for example, the displayed text when entering the recording mode can be the first two sentences in the prompt text data); the client can collect the user initial voice of the target user, transmit the user initial voice to the background server of the video application, and send a text conversion instruction to the background server. After the background server receives the user initial voice and the instruction sent by the client, it can execute step S15, detect the user initial voice by voice activity detection technology (VAD technology), delete noise and silence in the user initial voice, and obtain the user voice corresponding to the target user (that is, valid voice data). It should be noted that step S15 can be executed by the client through a local voice activity detection module, or the background server can execute it by using VAD technology.

[0152] Step S16, the background server can use a fast text conversion model to convert the user voice into text (i.e. user voice text), continue to step S17, convert the user voice text into pinyin (in the embodiment of the application, the default text prompt data is Chinese), and then step S18 can be executed. The background server can obtain the prompt text data input by the target user, convert the prompt text data into pinyin, match the pinyin of the user voice text with the pinyin of the prompt text data, continue to step S19, find the position of the text in the prompt text data that matches the user voice, and transmit the position of the text in the prompt text data to the client.

[0153] Step S20, after receiving the text position transmitted by the background server, the client can determine the target text corresponding to the user voice according to the text position, and highlight the target text in the recording page of the client, i.e. the prompt text data can be displayed by scrolling according to the text position; when the target user reads the last character in the prompt text data, the client can execute step S21 to end the video recording service. Of course, the target user can trigger the complete recording control in the recording page or trigger the cancel recording control in the recording page to end the video recording service.

[0154] After ending the video recording service, the client can transmit the recording video corresponding to the video recording service (i.e. the above-mentioned original video data) to the background server, and send a text conversion instruction to the background server. After receiving the text conversion instruction, the background server can execute step S22 to convert the voice data contained in the recording video into text (i.e. target text result) using an accurate text conversion model, and obtain the time when the text appears in the recording video, which can also be called the time stamp of the text in the recording video. At this time, the background server can execute steps S23 and S24 in parallel.

[0155] Step S23, the background server can compare the target text result with the prompt text data to find the error part (i.e. the voice data corresponding to the error text) in the recorded video; step S24, the background server can find the pause part in the user voice contained in the recorded video through the time (i.e. the time stamp) when the text appears in the recorded video. The background server can transmit the error part and the pause part in the recorded video to the client. After receiving the error part and the pause part transmitted by the background server, the client can perform step S25 to provide different clipping optimization modes for the target user in the client according to the error part and the pause part. The target user can select a suitable clipping optimization mode from the multiple clipping optimization modes provided by the client, and the client can clip and optimize the recorded video based on the clipping optimization mode selected by the target user to obtain the final target video data.

[0156] In the embodiments of the present application, after the user inputs the prompt text data in the video application, the user can start the video recording service through voice, and the video recording service provides a teleprompter function for the user during recording. The target text matching the user's voice can be located in the prompt text data, and the target text can be highlighted in the video application. The target text displayed in the video application matches the content of the user's speech, which can improve the effectiveness of the text prompt function of the video recording service, reduce the risk of recording failure due to forgetting words, and further improve the quality of the recorded video. Starting or stopping the video recording service through voice can reduce user operations in the video recording service and improve the effect of video recording. After the video recording service is completed, the recorded video in the video recording service can be automatically clipped and optimized, which can further improve the quality of the recorded video.

[0157] Please refer to Figure 12 , Figure 12 is a flowchart of a data processing method provided by the embodiments of the present application. It can be understood that the data processing method can be executed by a computer device, which can be a user terminal, or an independent server, or a cluster composed of multiple servers, or a system composed of a user terminal and a server, or a computer program application (including program code), which is not limited here. As shown in Figure 12 , the data processing method can include the following steps S201-S203:

[0158] Step S201, uploading prompt text data to a teleprompter application.

[0159] Specifically, the target user can input prompt text data in the teleprompter application or upload the edited prompt text data to the teleprompter application. The computer device can upload the prompt text data to the teleprompter application in response to the text input operation or the text upload operation of the target user, that is, when using the teleprompter function provided by the teleprompter application, the prompt text data needs to be uploaded to the teleprompter application. It should be noted that the computer device in the embodiments of the present application can be a device installed with a teleprompter application, which can also be referred to as a teleprompter.

[0160] In step S202, the user speech corresponding to the target user is collected, and the user speech text corresponding to the user speech is generated by text conversion.

[0161] Specifically, the computer device can collect the user initial speech of the target user, perform speech endpoint detection on the user initial speech, delete the noise and silence contained in the user initial speech, obtain the user speech (i.e., the valid speech data in the user initial speech) corresponding to the target user, and perform text conversion on the user speech to generate the user speech text corresponding to the user speech.

[0162] In step S203, the same text as the user speech text is determined as the target text in the prompt text data, and the target text is highlighted in the teleprompter application.

[0163] Specifically, the computer device can convert the user speech text into first syllable information, convert the prompt text data into second syllable information, compare the first syllable information with the second syllable information, determine the text position of the user speech text in the prompt text data, determine the target text matching the user speech in the prompt text data according to the text position, and highlight the target text in the teleprompter application. For more details of steps S202 and S203, please refer to the above Figure 3 The step S102 in the corresponding embodiment is not repeated here.

[0164] Optionally, the number of target users can be one or more, and different target users can correspond to different prompt text data; when the number of target users is one, the determination and display process of the target text in the teleprompter application can refer to the above Figure 3The step S102 in the corresponding embodiment; when the number of target users is multiple, after the computer device collects the user voice, the computer device can perform voiceprint recognition on the user voice, determine the user identity corresponding to the collected user voice according to the voiceprint recognition result, determine the target text corresponding to the user voice in the prompt text data corresponding to the user identity, and highlight the target text in the teleprompter application. The voiceprint recognition can refer to extracting the voiceprint features (for example, frequency spectrum, inverse frequency spectrum, formant, pitch, reflection coefficient, etc.) in the user voice data, and determining the user identity corresponding to the user voice by recognizing the voiceprint features, so the voiceprint recognition can also be called speaker recognition.

[0165] The number of target users is 2, that is, the target users include a first user and a second user, and the prompt text data includes a first prompt text corresponding to the first user and a second prompt text corresponding to the second user. The computer device can obtain the user voiceprint features in the user voice, determine the user identity corresponding to the user voice according to the user voiceprint features, if the user identity is the first user, determine the text same as the user voice text as the target text in the first prompt text, and highlight the target text in the teleprompter application, if the user identity is the second user, determine the text same as the user voice text as the target text in the second prompt text, and highlight the target text in the teleprompter application. In other words, when the number of target users is multiple, the user identity corresponding to the user voice needs to be determined first, and then the target text matching the user voice can be determined in the prompt text data corresponding to the user identity, and the target text can be highlighted, which can improve the effectiveness of the teleprompter function in the teleprompter application.

[0166] Please see Figure 13 , Figure 13 is an application scenario diagram of a teleprompter provided by the embodiment of the present application. The data processing process is described taking the teleprompter scene of a party as an example, such as Figure 13As shown, the script 90a (i.e., prompt text data) for the hosts during the gala can be pre-edited and uploaded to the teleprompter (which can be understood as the device where the teleprompter application is located, providing script prompts for the hosts). Script 90a can include the scripts of hosts A and B. After receiving script 90a, the teleprompter can save it locally. During the gala, the teleprompter can collect the voice data of all hosts in real time. When the teleprompter collects the user's voice, it can perform voiceprint recognition and determine the user's identity based on the voiceprint recognition results. When the user's identity in the collected voice is host A, the teleprompter can search for the target text matching the collected user's voice in host A's script (such as "With warm winter blessings, full of joy") and highlight "With warm winter blessings, full of joy" in the teleprompter.

[0167] When the user's voice is collected and the user's identity is Xiao B, the teleprompter can search for the target text that matches the collected user's voice in the host Xiao B's script (such as "In the past year, we have worked hard"), and highlight "In the past year, we have worked hard" in the teleprompter.

[0168] In this embodiment, the teleprompter can highlight the sentence that the target user is reading aloud, and automatically recognize the target user's voice as the target user reads aloud. The prompt text data is scrolled in the teleprompter, which can improve the effectiveness of the text prompt function in the teleprompter.

[0169] Please see Figure 14 , Figure 14 This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application. This data processing apparatus can perform the above-described... Figure 3 The steps in the corresponding embodiments, such as Figure 14 As shown, the data processing device 1 may include: a startup module 101, a display module 102, and an acquisition module 103;

[0170] The startup module 101 is used to respond to the service startup operation in the video application and start the video recording service in the video application.

[0171] The display module 102 is used to collect user voice during video recording, determine the target text associated with the user voice in the prompt text data associated with the video recording, and highlight the target text.

[0172] The acquisition module 103 is used to acquire the target video data corresponding to the video recording service when the target text is located at the end of the prompt text data.

[0173] The specific function implementation manners of the starting module 101, the display module 102, and the acquisition module 103 can be referred to the above Figure 3 The steps S101-S103 in the corresponding embodiments will not be repeated here.

[0174] In some possible implementation manners, the data processing apparatus 1 can further include a first recording page display module 104, an editing module 105, a first estimated duration display module 106, a second recording page display module 107, a text uploading module 108, and a second estimated duration display module 109.

[0175] The first recording page display module 104 is configured to display a recording page in the video application in response to a triggering operation on a teleprompter shooting entrance in the video application. The recording page includes a text input area.

[0176] The editing module 105 is configured to display prompt text data determined by an information editing operation on the text input area in the text input area in response to the information editing operation.

[0177] The first estimated duration display module 106 is configured to display the number of prompt texts and a video estimated duration corresponding to the prompt text data in the text input area when the number of prompt texts corresponding to the prompt text data is greater than a number threshold.

[0178] The second recording page display module 107 is configured to display a recording page in the video application in response to a triggering operation on a teleprompter shooting entrance in the video application. The recording page includes a text uploading control and a text input area.

[0179] The text uploading module 108 is configured to determine text content uploaded to the recording page as prompt text data in response to a triggering operation on the text uploading control, and display the prompt text data in the text input area.

[0180] The second estimated duration display module 109 is configured to display the number of prompt texts corresponding to the prompt text data and a video estimated duration corresponding to the prompt text data.

[0181] The specific function implementation manners of the first recording page display module 104, the editing module 105, the first estimated duration display module 106, the second recording page display module 107, the text uploading module 108, and the second estimated duration display module 109 can be referred to the above Figure 3The step S101 in the corresponding embodiment, here will not be described. Among them, when the first recording page display module 104, the editing module 105, the first estimated duration display module 106 execute corresponding operation, the second recording page display module 107, the text upload module 108, the second estimated duration display module 109 all suspend the operation of executing; when the second recording page display module 107, the text upload module 108, the second estimated duration display module 109 execute corresponding operation, the first recording page display module 104, the editing module 105, the first estimated duration display module 106 all suspend the operation of executing. Among them, the first recording page display module 104 and the second recording page display module 107 can be combined into the same recording page display module; the first estimated duration display module 106 and the second estimated duration display module 109 can be combined into the same estimated duration display module.

[0182] In some possible implementation manners, the service starting operation includes a voice starting operation.

[0183] The starting module 101 can include: a countdown animation display unit 1011, a recording service starting unit 1012.

[0184] The countdown animation display unit 1011 is configured to display a recording countdown animation associated with a video recording service in a recording page of the video application in response to a voice starting operation in the video application.

[0185] The recording service starting unit 1012 is configured to start and execute the video recording service in the video application when the recording countdown animation ends.

[0186] The specific function implementation manners of the countdown animation display unit 1011 and the recording service starting unit 1012 can be referred to the above Figure 3 The step S101 in the corresponding embodiment, here will not be described.

[0187] In some possible implementation manners, the recording countdown animation includes an animation cancel control.

[0188] The data processing apparatus 1 can further include: a countdown animation cancel module 110.

[0189] The countdown animation cancel module 110 is configured to cancel displaying the recording countdown animation and start and execute the video recording service in the video application in response to a triggering operation on the animation cancel control.

[0190] The specific function implementation manner of the countdown animation cancel module 110 can be referred to the above Figure 3 The step S101 in the corresponding embodiment, here will not be described.

[0191] In some possible implementation manners, the display module 102 can include: a voice endpoint detection unit 1021, a target text determination unit 1022, and a target text display unit 1023.

[0192] The voice endpoint detection unit 1021 is configured to collect initial voice of a user in a video recording service, perform voice endpoint detection on the initial voice of the user, and determine valid voice data in the initial voice of the user as user voice.

[0193] The target text determination unit 1022 is configured to convert the user voice into user voice text, perform text matching on the user voice text and prompt text data associated with the video recording service, and determine target text matched with the user voice text in the prompt text data.

[0194] The target text display unit 1023 is configured to highlight the target text in a recording page of the video recording service.

[0195] Specific function implementation manners of the voice endpoint detection unit 1021, the target text determination unit 1022, and the target text display unit 1023 can refer to step S102 in the above-mentioned corresponding embodiments, and will not be described here again. Figure 3

[0196] In some possible implementation manners, the target text determination unit 1022 can include: a syllable information acquisition subunit 10221 and a syllable matching subunit 10222.

[0197] The syllable information acquisition subunit 10221 is configured to acquire first syllable information corresponding to the user voice text, and acquire second syllable information corresponding to the prompt text data associated with the video recording service.

[0198] The syllable matching subunit 10222 is configured to acquire target syllable information same as the first syllable information in the second syllable information, and determine target text corresponding to the target syllable information in the prompt text data.

[0199] Specific function implementation manners of the syllable information acquisition subunit 10221 and the syllable matching subunit 10222 can refer to step S102 in the above-mentioned corresponding embodiments, and will not be described here again. Figure 3

[0200] In some possible implementation manners, the target text display unit 1023 can include: a prompt region determination subunit 10231 and a highlighting subunit 10232.

[0201] The prompt region determination subunit 10231 is configured to determine a text prompt region corresponding to the target text in a recording page of the video recording service. ​​

[0202] The highlighting sub-unit 10232 is configured to highlight the target text in the text prompt area according to the text position of the target text in the prompt text data.

[0203] The specific implementation of the prompt area determining sub-unit 10231 and the highlighting sub-unit 10232 can refer to the step S102 in the above Figure 3 corresponding embodiments, and will not be described here again.

[0204] In some possible implementation manners, the recording page includes a cancel recording control;

[0205] The data processing apparatus 1 can further include a recording cancel module 111, a recording prompt information display module 112, and a re-recording module 113.

[0206] The recording cancel module 111 is configured to cancel the video recording service and delete the video data recorded by the video recording service in response to a triggering operation on the cancel recording control.

[0207] The recording prompt information display module 112 is configured to generate recording prompt information for the video recording service and display the recording prompt information in the recording page. The recording prompt information includes a re-recording control.

[0208] The re-recording module 113 is configured to switch the display of the target text in the recording page to the prompt text data in response to a triggering operation on the re-recording control.

[0209] The specific implementation of the recording cancel module 111, the recording prompt information display module 112, and the re-recording module 113 can refer to the step S102 in the above Figure 3 corresponding embodiments, and will not be described here again.

[0210] In some possible implementation manners, the recording page includes a complete recording control;

[0211] The data processing apparatus 1 can include a recording complete module 114.

[0212] The recording complete module 114 is configured to stop the video recording service and determine the video data recorded by the video recording service as target video data in response to a triggering operation on the complete recording control.

[0213] The specific implementation of the recording complete module 114 can refer to the step S102 in the above Figure 3 corresponding embodiments, and will not be described here again.

[0214] In some possible implementation manners, the acquisition module 103 can include: an original video acquisition unit 1031, an optimization control display unit 1032, an optimization mode display unit 1033, and an optimization processing unit 1034.

[0215] The original video acquisition unit 1031 is configured to stop the video recording service when the text position of the target text in the prompt text data is the end position in the prompt text data, and determine video data recorded by the video recording service as original video data.

[0216] The optimization control display unit 1032 is configured to display the original video data and a clip optimization control corresponding to the original video data in a clip page of the video application.

[0217] The optimization mode display unit 1033 is configured to display M clip optimization modes for the original video data in response to a triggering operation on the clip optimization control. M is a positive integer.

[0218] The optimization processing unit 1034 is configured to perform clip optimization processing on the original video data according to a clip optimization mode determined by a selection operation on the M clip optimization modes, to obtain target video data corresponding to the video recording service.

[0219] The specific function implementation manners of the original video acquisition unit 1031, the optimization control display unit 1032, the optimization mode display unit 1033, and the optimization processing unit 1034 can be refer to step S103 in the above-mentioned Figure 3 corresponding embodiments, which will not be described herein again.

[0220] In some possible implementation manners, the optimization processing unit 1034 can include: a first speech conversion subunit 10341, a text comparison subunit 10342, a speech deletion subunit 10343, a second speech conversion subunit 10344, a timestamp acquisition subunit 10345, and a speech pause segment determination subunit 10346.

[0221] The first speech conversion subunit 10341 is configured to, if the clip optimization mode determined by the selection operation is a first clip mode, acquire target speech data contained in the original video data, and convert the target speech data into a target text result.

[0222] The text comparison subunit 10342 is configured to compare the target text result with the prompt text data, and determine text that is different from the prompt text data in the target text result as error text.

[0223] The speech deletion subunit 10343 is configured to delete speech data corresponding to the error text in the original video data, to obtain the target video data corresponding to the video recording service.

[0224] The second speech conversion subunit 10344 is configured to convert target speech data contained in the original video data into target text results if the selected operation-determined clipping optimization mode is the second clipping mode, and determine text in the target text results that is different from the prompt text data as error text;

[0225] The timestamp acquisition subunit 10345 is configured to divide the target text results into N text characters, and acquire timestamps of the N text characters in the target speech data; N is a positive integer.

[0226] The speech pause segment determination subunit 10346 is configured to determine speech pause segments in the target speech data according to the timestamps, delete speech data corresponding to the speech pause segments and the error text in the original video data, and obtain target video data corresponding to the video recording service.

[0227] The specific function implementation manners of the first speech conversion subunit 10341, the text comparison subunit 10342, the speech deletion subunit 10343, the second speech conversion subunit 10344, the timestamp acquisition subunit 10345, and the speech pause segment determination subunit 10346 can be referred to the step S103 in the above-mentioned Figure 3 corresponding embodiments, which will not be described herein again. When the first speech conversion subunit 10341, the text comparison subunit 10342, and the speech deletion subunit 10343 perform corresponding operations, the second speech conversion subunit 10344, the timestamp acquisition subunit 10345, and the speech pause segment determination subunit 10346 suspend the execution of operations; when the second speech conversion subunit 10344, the timestamp acquisition subunit 10345, and the speech pause segment determination subunit 10346 perform corresponding operations, the first speech conversion subunit 10341, the text comparison subunit 10342, and the speech deletion subunit 10343 suspend the execution of operations.

[0228] In some possible implementation manners, the data processing apparatus 1 can further include a user speech speed determination module 115 and a speech speed prompt information display module 116.

[0229] The user speech speed determination module 115 is configured to acquire a speech duration corresponding to user initial speech and a number of speech characters contained in the user initial speech, and determine a ratio of the number of speech characters to the speech duration as a user speech speed.

[0230] The speech speed prompt information display module 116 is configured to display speech speed prompt information in a recording page when the user speech speed is greater than a speech speed threshold value; the speech speed prompt information is used to prompt a target user associated with the video recording service to reduce the user speech speed.

[0231] The specific function implementation manners of the user speech speed determination module 115 and the speech speed prompt information display module 116 can be referred to the above Figure 3 The step S102 in the corresponding embodiment will not be repeated here.

[0232] In some possible implementation manners, the error text includes K error subtexts, K being a positive integer;

[0233] The data processing apparatus 1 can further include an error frequency determination module 117, an error type identification module 118, and a tutorial video pushing module 119.

[0234] The error frequency determination module 117 is configured to determine an error frequency in the video recording service according to the K error subtexts and a video duration corresponding to the original video data.

[0235] The error type identification module 118 is configured to identify a speech error type corresponding to each of the K error subtexts when the error frequency is greater than an error threshold.

[0236] The tutorial video pushing module 119 is configured to push, in the video application, a tutorial video associated with the speech error type to a target user associated with the video recording service.

[0237] The specific function implementation manners of the error frequency determination module 117, the error type identification module 118, and the tutorial video pushing module 119 can be referred to the above Figure 15 The step S103 in the corresponding embodiment will not be repeated here.

[0238] In the embodiments of the present application, after the user inputs the prompt text data in the video application, the user can start the video recording service through voice, and the user is provided with a teleprompter function in the recording process of the video recording service. The target text matched with the user voice can be located in the prompt text data, and the target text can be highlighted in the video application, that is, the target text displayed in the video application is matched with the content being delivered by the user, which can improve the effectiveness of the text prompt function of the video recording service, reduce the risk of recording failure caused by the user forgetting the words, and further improve the quality of the recorded video. The video recording service can be started or stopped through the user voice, which can reduce the user operation in the video recording service and improve the effect of video recording. After the video recording service ends, the recorded video in the video recording service can be automatically edited and optimized, which can further improve the quality of the recorded video.

[0239] Please refer to Figure 15 , Figure 12 The structure schematic diagram of the data processing apparatus provided in the embodiments of the present application is implemented. The data processing apparatus can execute the above Figure 15The steps in the corresponding embodiments, such as Figure 12 As shown in the corresponding embodiments, the data processing apparatus 2 can include: a prompt text uploading module 21, a user voice collecting module 22, and a user voice text display module 23.

[0240] The prompt text uploading module 21 is configured to upload the prompt text data to the teleprompter application.

[0241] The user voice collecting module 22 is configured to collect the user voice corresponding to the target user, perform text conversion on the user voice, and generate user voice text corresponding to the user voice.

[0242] The user voice text display module 23 is configured to determine the same text as the user voice text as the target text in the prompt text data, and highlight the target text in the teleprompter application.

[0243] The specific implementation of the prompt text uploading module 21, the user voice collecting module 22, and the user voice text display module 23 can be referred to the steps S201-S203 in the above-mentioned Figure 12 corresponding embodiments, which will not be repeated here.

[0244] The target user includes a first user and a second user, and the prompt text data includes first prompt text corresponding to the first user and second prompt text corresponding to the second user.

[0245] The user voice text display module 23 can include: a user identity determining unit 231, a first determining unit 232, and a second determining unit 233.

[0246] The user identity determining unit 231 is configured to obtain the user voiceprint feature in the user voice, and determine the user identity corresponding to the user voice according to the user voiceprint feature.

[0247] The first determining unit 232 is configured to, if the user identity is the first user, determine the same text as the user voice text as the target text in the first prompt text, and highlight the target text in the teleprompter application.

[0248] The second determining unit 233 is configured to, if the user identity is the second user, determine the same text as the user voice text as the target text in the second prompt text, and highlight the target text in the teleprompter application.

[0249] The specific implementation of the user identity determining unit 231, the first determining unit 232, and the second determining unit 233 can be referred to the step S203 in the above-mentioned Figure 16 corresponding embodiments, which will not be repeated here.

[0250] In the embodiment of the present application, the teleprompter can highlight the sentence being read by the target user and automatically identify the voice of the target user as the target user reads. The prompt text data is displayed in the teleprompter, which can improve the effectiveness of the text prompt function of the teleprompter.

[0251] Please refer to Figure 16 , Figure 16 is a structural schematic diagram of a computer device provided in an embodiment of the present application. As shown in Figure 16 , the computer device 1000 can include a processor 1001, a network interface 1004 and a memory 1005. In addition, the computer device 1000 can further include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between the components. The user interface 1003 can include a display, a keyboard, and optionally the user interface 1003 can further include a standard wired interface and a wireless interface. Optionally, the network interface 1004 can include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 can be a high-speed RAM memory or a non-volatile memory such as at least one disk memory. Optionally, the memory 1005 can also be at least one storage device located away from the aforementioned processor 1001. As shown in Figure 16 , the memory 1005 as a computer readable storage medium can include an operating system, a network communication module, a user interface module and a device control application.

[0252] In the computer device 1000 as shown in Figure 3 , the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to realize:

[0253] starting a video recording service in the video application in response to a service starting operation in the video application;

[0254] collecting user voice in the video recording service, determining target text associated with the user voice in the prompt text data associated with the video recording service, and highlighting the target text;

[0255] when the text position of the target text in the prompt text data is the end position of the prompt text data, obtaining target video data corresponding to the video recording service.

[0256] It should be understood that the computer device 1000 described in the embodiments of the present application can perform the foregoingFigure 14 The description of the data processing method in the corresponding embodiment can also be performed by the data processing device 1 described in the foregoing Figure 17 The description of the data processing device 1 in the corresponding embodiment will not be repeated here. In addition, the beneficial effects of using the same method will not be repeated.

[0257] Please refer to Figure 17 , Figure 17 is a structural schematic diagram of a computer device provided by an embodiment of the present application. As shown in Figure 17 , the computer device 2000 can include a processor 2001, a network interface 2004, and a memory 2005, in addition to which the computer device 2000 can further include a user interface 2003 and at least one communication bus 2002. The communication bus 2002 is used to realize the connection and communication between the components. The user interface 2003 can include a display, a keyboard, and can optionally include a standard wired interface, a wireless interface. Optionally, the network interface 2004 can include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 2005 can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory. Optionally, the memory 2005 can also be at least one storage device located away from the aforementioned processor 2001. As shown in Figure 17 , the memory 2005, as a computer readable storage medium, can include an operating system, a network communication module, a user interface module, and a device control application program.

[0258] In the computer device 2000 as shown in Figure 6 , the network interface 2004 can provide network communication functions; the user interface 2003 is mainly used to provide an input interface for the user; and the processor 2001 can be used to call the device control application program stored in the memory 2005 to realize:

[0259] uploading the prompt text data to a teleprompter application;

[0260] collecting user speech corresponding to a target user, performing text conversion on the user speech to generate user speech text corresponding to the user speech;

[0261] in the prompt text data, determining the same text as the user speech text as a target text, and highlighting the target text in the teleprompter application.

[0262] It should be understood that the computer device 2000 described in the embodiments of the present application can perform the description of the data processing method in the foregoing Figure 14 corresponding embodiment, and also can perform the description of the data processing method in the foregoingFigure 3 The description of the data processing device 2 in any of the corresponding embodiments will not be repeated here. In addition, the description of the beneficial effects of using the same method will also not be repeated.

[0263] In addition, it should be noted that the embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores the computer program executed by the data processing device 1 mentioned above, and the computer program includes program instructions, which can execute the above-mentioned Figure 11 、 Figure 12 and Figure 3 The description of the data processing method in any of the corresponding embodiments will not be repeated here. In addition, the description of the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer readable storage medium embodiments of the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on one computing device for execution, or on multiple computing devices located in one place for execution, or on multiple computing devices distributed in multiple places and interconnected through a communication network for execution. The multiple computing devices distributed in multiple places and interconnected through a communication network can constitute a blockchain system.

[0264] In addition, it should be noted that the embodiments of the present application also provide a computer program product or computer program, which can include computer instructions that can be stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor can execute the computer instructions to make the computer device execute the above-mentioned Figure 11 、 Figure 12 and ​ The description of the data processing method in any of the corresponding embodiments will not be repeated here. In addition, the description of the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program product or computer program embodiments of the present application, please refer to the description of the method embodiments of the present application.

[0265] It should be noted that for each of the above method embodiments, in order to simply describe, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0266] The steps in the method embodiments of the present application can be adjusted in sequence, combined and reduced according to actual needs.

[0267] The modules in the device embodiments of the present application can be combined, divided and reduced according to actual needs.

[0268] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments can be included. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM).

[0269] The above only discloses the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope of the present application.

Claims

1. A data processing method, characterized by, The method comprises: starting a video recording service in the video application in response to a service starting operation in the video application; collecting user speech in the video recording service, converting the user speech into user speech text by using a fast speech-to-text model, determining target text corresponding to the user speech in prompt text data associated with the video recording service according to pronunciation of the user speech text and pronunciation of the prompt text data, and highlighting the target text; wherein the fast speech-to-text model only considers that pronunciation of converted text is consistent with the user speech when converting. When a text position of the target text in the prompt text data is an end position in the prompt text data, obtaining target video data corresponding to the video recording service.

2. The method of claim 1, wherein, The method further comprises: displaying a recording page in the video application in response to a trigger operation on a teleprompter shooting entrance in the video application; the recording page comprises a text input area; displaying prompt text data determined by an information editing operation on the text input area in the text input area in response to the information editing operation on the text input area; when a number of prompt characters corresponding to the prompt text data is greater than a number threshold, displaying the number of prompt characters and a video estimated duration corresponding to the prompt text data in the text input area.

3. The method of claim 1, wherein, The method further comprises: displaying a recording page in the video application in response to a trigger operation on a teleprompter shooting entrance in the video application; the recording page comprises a text upload control and a text input area; determining text content uploaded to the recording page as prompt text data in response to a trigger operation on the text upload control, and displaying the prompt text data in the text input area; displaying a number of prompt characters corresponding to the prompt text data and a video estimated duration corresponding to the prompt text data.

4. The method of claim 1, wherein, The service starting operation comprises a voice starting operation; The method further comprises: displaying a recording countdown animation associated with the video recording service in a recording page of the video application in response to a voice starting operation in the video application; starting and executing the video recording service in the video application when the recording countdown animation ends.

5. The method of claim 4, wherein, The recording countdown animation comprises an animation cancel control. The method further comprises: canceling display of the recording countdown animation and starting and executing the video recording service in the video application in response to a trigger operation on the animation cancel control.

6. The method of claim 1, wherein, Before converting the user speech into user speech text, the method further comprises: collecting user initial speech in the video recording service, performing speech endpoint detection on the user initial speech, and determining valid speech data in the user initial speech as the user speech.

7. The method of claim 1, wherein, The highlighting of the target text comprises: highlighting the target text in a recording page of the video recording service.

8. The method of claim 1, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:

9. The method of claim 7, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:

10. The method of claim 9, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:

11. The method of claim 9, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:

12. The method of claim 1, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:

13. The method of claim 12, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the text comparison is performed between the target text result and the prompt text data, and text in the target text result that is different from the prompt text data is determined as error text; speech data corresponding to the error text is deleted from the original video data, to obtain target video data corresponding to the video recording service.

14. The method of claim 12, wherein, The clipping optimization manner determined according to the selection operation is used to perform clipping optimization processing on the original video data, to obtain target video data corresponding to the video recording service, including: If the clipping optimization manner determined by the selection operation is a second clipping manner, target speech data included in the original video data is converted into a target text result, and text in the target text result that is different from the prompt text data is determined as error text; The target text result is divided into N text characters, and timestamps of the N text characters in the target speech data are obtained; N is a positive integer; According to the timestamps, a speech pause segment in the target speech data is determined, and speech data corresponding to the speech pause segment and the error text is deleted from the original video data, to obtain target video data corresponding to the video recording service.

15. The method of claim 6, wherein, Further comprising: obtaining a speech duration corresponding to the user initial speech and a number of speech characters included in the user initial speech, and determining a user speech speed as a ratio of the number of speech characters to the speech duration; when the user speech speed is greater than a speech speed threshold, displaying a speech speed prompt information in the recording page; The speech speed prompt information is used to prompt a target user associated with the video recording service to reduce the user speech speed.

16. The method according to any of claims 13-14, characterized by, The error text includes K error subtexts, and K is a positive integer; The method further comprises: determining an error frequency in the video recording service according to the K error subtexts and a video duration corresponding to the original video data; when the error frequency is greater than an error threshold, identifying a speech error type corresponding to each of the K error subtexts; pushing a tutorial video associated with the speech error type to a target user associated with the video recording service in the video application.

17. A data processing method, characterized by, comprising: uploading prompt text data to a teleprompter application; collecting user speech corresponding to a target user, and using a fast speech-to-text model to perform text conversion on the user speech to generate user speech text corresponding to the user speech; wherein the fast speech-to-text model only considers that the pronunciation of the converted text remains consistent with the user speech during conversion; determining target text corresponding to the user speech in the prompt text data according to the pronunciation of the user speech text and the pronunciation of the prompt text data, and highlighting the target text in the teleprompter application.

18. The method of claim 17, wherein, The target user includes a first user and a second user, and the prompt text data includes first prompt text corresponding to the first user and second prompt text corresponding to the second user; The step of determining the target text corresponding to the user's voice in the prompt text data based on the pronunciation of the user's voice text and the pronunciation of the prompt text data, and highlighting the target text in the prompting application, includes: Obtain the user's voiceprint features from the user's voice, and determine the user's identity corresponding to the user's voice based on the user's voiceprint features; If the user is the first user, then based on the pronunciation of the user's voice text and the pronunciation of the prompt text data, the target text corresponding to the user's voice is determined in the first prompt text, and the target text is highlighted in the prompting application; If the user is the second user, then based on the pronunciation of the user's voice text and the pronunciation of the prompt text data, the target text corresponding to the user's voice is determined in the second prompt text, and the target text is highlighted in the prompting application.

19. A data processing apparatus, characterized by include: The startup module is used to respond to the service startup operation in the video application and start the video recording service in the video application; The display module is used to collect user voice during the video recording service, and uses a fast speech-to-text model to convert the user voice into user voice-text. Based on the pronunciation of the user voice-text and the pronunciation of the prompt text data associated with the video recording service, the target text corresponding to the user voice is determined in the prompt text data, and the target text is highlighted. The fast speech-to-text model only considers that the pronunciation of the converted text is consistent with the user voice during the conversion. The acquisition module is used to acquire the target video data corresponding to the video recording service when the target text is located at the end of the prompt text data.

20. A computer device, comprising: Including memory and processor; The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to invoke the computer program so that the computer device performs the method according to any one of claims 1 to 16, or performs the method according to any one of claims 17 to 18.

21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1 to 16, or to perform the method of any one of claims 17 to 18.

22. A computer program product, characterised in that, The computer program product includes computer instructions that, when executed by a processor, implement the method of any one of claims 1 to 16, or implement the method of any one of claims 17 to 18.

Citation Information

Patent Citations

  • Multimedia data recording method and device and electronic equipment

    CN111372119A

  • Audio and video correction method and device, medium and computing equipment

    CN111885313A