Model training method and device, video generation method and device, equipment and medium
By conducting multiple training and optimization of the segmentation model, the problem of inefficient and low accuracy in processing training videos in the prior art is solved, and a higher model accuracy and video editing effect is achieved.
Patent Information
- Application Number
- CN202411717154.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-16
AI Technical Summary
The existing segmentation model is inefficient when processing training videos, and the outputted cutting points are very different from the segmentation points that users actually need, so the model accuracy is not high.
A model training method is provided, by obtaining the first segmentation model and the training data set, training the first segmentation model according to the reference segmentation point to generate the second segmentation model, and constructing the first scoring model based on the second segmentation model, and finally obtaining the target segmentation model through multiple training and optimization.
The accuracy of the model is improved, the accuracy of the theme recognition of video content and the filming effect of video editing are improved, so that the cutting points output when the model processes the video to be edited are closer to the segmentation points that the user actually needs.
Smart Images

Figure CN120014503A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a model training method, video generation method, device, equipment and medium. Background Art
[0002] When editing training videos, the industry generally uses software for manual editing. Specifically, when editing, the training video needs to be segmented according to the subject content to obtain the required target video. Therefore, at present, some solutions use segmentation models to edit training videos;
[0003] However, most existing segmentation models use shot segmentation to segment and edit videos, but this method is very inefficient for training videos, because training videos are generally dominated by long shots, and there are not many switches in the shots. As a result, the output cutting points when processing the video are quite different from the segmentation points actually required by the user, and the model accuracy is not high. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to provide a model training method, a video generation method, an apparatus, a device and a medium, aiming to obtain a target segmentation model for segmenting the video to be edited input into the model based on the content theme of the sample video.
[0005] In a first aspect, an embodiment of the present application provides a model training method, comprising:
[0006] Obtaining a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points;
[0007] Training the first segmentation model according to the reference segmentation point to generate a second segmentation model;
[0008] Building a first scoring model based on the second segmentation model, the first scoring model is used to assign segmentation scores to sentences in the sample video, the segmentation scores representing the distance from the sentence to the model segmentation point;
[0009] Training the first scoring model according to the reference segmentation point to generate a second scoring model;
[0010] Inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence;
[0011] Constructing an objective function, the output of which matches the expected value of the segmentation score and the update step size of the second segmentation model;
[0012] The second segmentation model is trained according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
[0013] In some implementations, the output of the objective function is proportional to the expected value of the segmentation score and inversely proportional to the update step size of the second segmentation model, and the second segmentation model is trained according to the sentence data set until the output of the objective function converges, including:
[0014] The second segmentation model is copied to generate a third segmentation model, and the update step size is the strategy difference between the segmentation points of the model generated by the second segmentation model and the segmentation points of the model generated by the third segmentation model;
[0015] Optimizing the strategy of generating model segmentation points by the third segmentation model according to the sentence data set;
[0016] The model parameters of the third segmentation model are copied to the second segmentation model, so that the second segmentation model adopts the same strategy of generating model segmentation points as the third segmentation model, until the output of the objective function converges.
[0017] In some implementations, training the first segmentation model according to the reference segmentation point to generate the second segmentation model includes:
[0018] Establishing a first loss function between the reference segmentation point and the model segmentation point, the first loss function is used to characterize the error of the model segmentation point relative to the reference segmentation point;
[0019] The sample video is input into the first segmentation model so that the first segmentation model generates model segmentation points, and the parameters of the first segmentation model are optimized according to the first loss function until the first loss function converges, thereby obtaining a second segmentation model.
[0020] In some embodiments, the second segmentation model includes a vector layer and a segmentation layer, the vector layer is used to obtain a feature vector of a sentence in a sample video, and the segmentation layer is used to generate a model segmentation point according to the feature vector of the sentence;
[0021] Building a first scoring model based on the second segmentation model includes:
[0022] Extracting a vector layer of a second segmentation model;
[0023] Construct a fully connected layer, which is used to generate the segmentation score of the corresponding sentence based on the input feature vector;
[0024] The first scoring model is constructed based on the vector layer and the fully connected layer.
[0025] In some implementations, training the first score model according to the reference segmentation point to generate the second score model includes:
[0026] Assign corresponding reference scores to sentences in the sample video according to the reference segmentation points marked in the sample video, where the reference scores represent the distance from the sentences to the reference segmentation points;
[0027] Establishing a second loss function between the reference score and the segmentation score, the second loss function is used to characterize the error of the segmentation score of the sentence relative to the reference score;
[0028] The sample video is input into the first scoring model so that the first scoring model generates a segmentation score, and the parameters of the second scoring model are optimized according to the second loss function until the second loss function converges to obtain the second scoring model.
[0029] In some implementations, inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence includes:
[0030] Inputting the sample video into at least one of a second score model and a second segmentation model to generate a feature vector of the sentence;
[0031] Inputting the sample video into a second scoring model to generate a segmentation score for the sentence;
[0032] Inputting the sample video into a second segmentation model to generate model segmentation points of the sample video;
[0033] The feature vector, segmentation score, and information about whether the sentence is marked with a model segmentation point are integrated to generate a sentence dataset corresponding to the sentence.
[0034] In a second aspect, an embodiment of the present application further provides a video generation method, the method comprising:
[0035] Receive the video to be edited and the editing instruction;
[0036] Calling a target segmentation model according to the clipping instruction, wherein the target segmentation model is trained by using any one of the model training methods provided in the embodiments of the present application;
[0037] The video to be edited is input into the target segmentation model so that the target segmentation model generates model segmentation points corresponding to the video to be edited, and the video to be edited is segmented according to the model segmentation points to generate at least one target video.
[0038] In a third aspect, the embodiment of the present application further provides a model training device, comprising:
[0039] A sample acquisition module, used to acquire a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points;
[0040] A first training module, used for training the first segmentation model according to the reference segmentation point to generate a second segmentation model;
[0041] A scoring model module, used to construct a first scoring model based on the second segmentation model, the first scoring model is used to assign segmentation scores to sentences in the sample video, and the segmentation scores represent the distance from the sentence to the model segmentation point;
[0042] A second training module, used for training the first scoring model according to the reference segmentation point to generate a second scoring model;
[0043] A sample processing module, used to input the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence;
[0044] A function building module, used to build an objective function, the output of which matches the expected value of the segmentation score and the learning rate of the second segmentation model;
[0045] The third training module is used to train the second segmentation model according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
[0046] In a fourth aspect, an embodiment of the present application further provides a computer device, the computer device comprising a memory and a processor;
[0047] Memory for storing computer programs;
[0048] A processor is used to execute a computer program and implement the steps of any model training method provided in the embodiments of the present application when executing the computer program.
[0049] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the steps of any model training method provided in the embodiment of the present application.
[0050] In summary, the embodiments of the present application provide a model training method, a video generation method, an apparatus, a device and a medium, wherein the method comprises: obtaining a first segmentation model and a training data set, wherein the training data set comprises a sample video and a reference segmentation point annotated in the sample video, the first segmentation model is used to generate a model segmentation point to segment the sample video according to the model segmentation point; the first segmentation model is trained according to the reference segmentation point to generate a second segmentation model; a first scoring model is constructed based on the second segmentation model, the first scoring model is used to assign a segmentation score to a sentence in the sample video, the segmentation score represents the distance from the sentence to the model segmentation point; the first scoring model is trained according to the reference segmentation point The method comprises the following steps: training to generate a second scoring model; inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to the sentences one by one; constructing an objective function, the output of which matches the expected value of the segmentation score and the update step size of the second segmentation model; training the second segmentation model according to the sentence data set until the output of the objective function converges, so as to obtain a target segmentation model for segmenting the video to be edited input into the model based on the content theme of the sample video, thereby improving the model accuracy, the accuracy of video content theme recognition and the film-forming effect of video editing, and making the cutting point output by the model when processing the video to be edited closer to the segmentation point actually required by the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 A flowchart of a model training method provided in an embodiment of the present application;
[0053] Figure 2 A schematic diagram of a flow chart of a step of training and generating a second segmentation model in a model training method provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of training a first segmentation model in a model training method provided in an embodiment of the present application;
[0055] Figure 4 A schematic flow chart of the step of constructing a first scoring model in a model training method provided in an embodiment of the present application;
[0056] Figure 5 A model structure diagram of a second segmentation model and a first scoring model in a model training method provided in an embodiment of the present application;
[0057] Figure 6 A schematic diagram of training a first scoring model in a model training method provided in an embodiment of the present application;
[0058] Figure 7 A schematic diagram of a flow chart of a video generation method provided in an embodiment of the present application;
[0059] Figure 8 A schematic block diagram of the structure of a model training device provided in an embodiment of the present application;
[0060] Fig. 9 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0062] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0063] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0064] When editing training videos, the industry generally uses software for manual editing. Specifically, when editing, the training video needs to be segmented according to the subject content to obtain the required target video. Therefore, at present, some solutions use segmentation models to edit training videos;
[0065] However, most existing segmentation models use shot segmentation to segment and edit videos, but this method is very inefficient for training videos, because training videos are generally dominated by long shots, and there are not many switches in the shots. As a result, the output cutting points when processing the video are quite different from the segmentation points actually required by the user, and the model accuracy is not high.
[0066] To solve the above problems, the embodiments of the present application provide a model training method, a video generation method, an apparatus, a device and a medium. The following is a detailed description of some embodiments of the present application in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0067] It should be noted that the model training method can be applied to a computer device, which can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, a wearable device, etc., or a server terminal, wherein the server terminal can be an independent server or a server cluster. In the embodiment of the present application, an independent server is used to perform the method as an example for specific description.
[0068] Please refer to Figure 1 , Figure 1 A flowchart of a model training method provided in an embodiment of the present application.
[0069] like Figure 1 As shown, the model training method includes steps S1 to S7.
[0070] Step S1: Obtain a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points.
[0071] Specifically, the server executing the method obtains a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video.
[0072] In some embodiments, the sample videos in the training data set are pre-stored in a server, and obtaining the training data set specifically includes: first calling the stored sample video and displaying the sample video to the user so that the user can mark the reference segmentation points in the displayed sample video; then generating the training data set based on the sample video and the reference segmentation points marked in the sample video.
[0073] Among them, the method of displaying the sample video to the user is, for example, outputting information carrying the sample video to the electronic terminal corresponding to the user, so that the user uses the electronic terminal to mark the reference segmentation points of the sample video, and outputting information carrying the sample video and the reference segmentation points marked in the sample video to the server through the electronic terminal.
[0074] For example, the reference segmentation points in the training data set are specifically marked according to the content theme in the sample video, so as to obtain two sub-videos with different content themes based on the reference segmentation points. Furthermore, the reference segmentation points in the training data set are specifically set between two sentences in the sample video.
[0075] Specifically, at least one reference segmentation point is marked in the sample video to segment the sample video into at least two sub-videos with different content themes. For example, if the sample video contains two adjacent sub-videos with content themes A and B, the reference segmentation point is marked between the two sub-videos.
[0076] It should also be noted that the first segmentation model is used to generate model segmentation points based on the input sample video to segment the sample video according to the model segmentation points. For example, the first segmentation model can be a GPT (Generative Pre-training Transformer) model.
[0077] The model training method provided in the embodiment of the present application is to perform model training on the basis of the first segmentation model to generate a target segmentation model. The target segmentation model segments the sample video based on the content theme of the sample video to improve the model accuracy, the accuracy of video content theme recognition and the film effect of video editing, so that the cutting point output when processing the video is closer to the segmentation point actually required by the user.
[0078] Step S2: training the first segmentation model according to the reference segmentation points to generate a second segmentation model.
[0079] Specifically, the first segmentation model generates a model segmentation point between two sentences in the sample video according to the input sample video, and there may be an error between the model segmentation point and the reference segmentation point, that is, the first segmentation model is not accurate enough, and the output cutting point is still different from the segmentation point actually required by the user. It should be noted that the error between the above-mentioned model segmentation point and the reference segmentation point specifically refers to the time difference between the model segmentation point and the reference segmentation point in the sample video.
[0080] See also Figure 2 and Figure 3 , Figure 2 A flowchart of the step of training and generating a second segmentation model in a model training method provided in an embodiment of the present application, Figure 3 A schematic diagram of training a first segmentation model in a model training method provided in an embodiment of the present application.
[0081] like Figure 2 and Figure 3 As shown, in some embodiments, step S2 trains the first segmentation model according to the reference segmentation point to generate a second segmentation model, including steps S21 to S22:
[0082] Step S21: establishing a first loss function between a reference segmentation point and a model segmentation point, wherein the first loss function is used to characterize an error of the model segmentation point relative to the reference segmentation point;
[0083] Step S22: Input the sample video into the first segmentation model so that the first segmentation model generates model segmentation points, and optimize the parameters of the first segmentation model according to the first loss function until the first loss function converges to obtain a second segmentation model.
[0084] It should be noted that the first loss function is used to characterize the error of the model segmentation point relative to the reference segmentation point, that is, the time difference between the model segmentation point and the reference segmentation point in the sample video. In the process of training the first segmentation model, the output of the first loss function changes accordingly, and the server also optimizes the model parameters of the first segmentation model according to the output of the instant first loss function. When the first loss function converges, it indicates that the output of the first loss function is stable and the training of the first segmentation model is completed.
[0085] For example, establishing the first loss function may be establishing a cross entropy loss function based on the time difference between the reference segmentation point and the model segmentation point in the sample video.
[0086] Specifically, training the first segmentation model to generate the second segmentation model is first to establish a first loss function between the reference segmentation point and the model segmentation point, then input the sample video into the first segmentation model so that the first segmentation model generates the model segmentation point, and the first loss function is updated according to the generated model segmentation point, after which the parameters of the first segmentation model are optimized according to the output of the first loss function to iteratively train the first segmentation model to obtain the second segmentation model.
[0087] After optimizing the parameters of the first segmentation model and under the condition that the first loss function has not converged, the server executing the present method will again input the sample video into the first segmentation model so that the first segmentation model generates model segmentation points, and update the first loss function according to the model segmentation points generated by the first segmentation model. Thereafter, the parameters of the first segmentation model are optimized according to the output of the first loss function to iteratively train the first segmentation model until the first loss function converges to obtain a second segmentation model.
[0088] Step S3: construct a first scoring model based on the second segmentation model, the first scoring model is used to assign segmentation scores to sentences in the sample video, and the segmentation scores represent the distance from the sentence to the model segmentation point.
[0089] After generating the second segmentation model, the server executing the method constructs a first scoring model based on the second segmentation model, and the first scoring model is used to assign segmentation scores to sentences in the sample video. It should be noted that the segmentation score represents the distance from the sentence to the model segmentation point closest to the sentence. The closer the distance from any sentence in the sample video to the model segmentation point, the higher the segmentation score assigned to the sentence.
[0090] For example, assume that the second segmentation model generates a model segmentation point at time T in the sample video, and the segmentation scores of the six sentences before and after time T in the sample video are specifically:
[0091] The segmentation score of the sixth sentence before T time is 0, the segmentation score of the fifth sentence before T time is 1, the segmentation score of the fourth sentence before T time is 2, the segmentation score of the third sentence before T time is 3, the segmentation score of the second sentence before T time is 4, the segmentation score of the first sentence before T time is 5, the segmentation score of the first sentence after T time is 5, the segmentation score of the second sentence after T time is 4, the segmentation score of the third sentence after T time is 3, the segmentation score of the fourth sentence after T time is 2, the segmentation score of the fifth sentence after T time is 1, and the segmentation score of the sixth sentence after T time is 0.
[0092] That is, the segmentation score assigned to any sentence in the sample video is inversely proportional to its distance to the nearest model segmentation point.
[0093] See also Figure 4 and Figure 5 , Figure 4 A flowchart of the step of constructing a first scoring model in a model training method provided in an embodiment of the present application is shown in FIG. Figure 5 A model structure diagram of a second segmentation model and a first scoring model in a model training method provided in an embodiment of the present application, wherein: Figure 5 (a) is the model structure diagram of the second segmentation model, and Figure 5 (b) is the model structure diagram of the first scoring model.
[0094] In some embodiments, the second segmentation model includes a vector layer and a segmentation layer, wherein the vector layer is used to obtain the feature vector of the sentence in the sample video, and the segmentation layer is used to generate model segmentation points according to the feature vector of the sentence. Similarly, the first segmentation model has the same model structure as the second segmentation model.
[0095] The first scoring model includes a vector layer and a fully connected layer, wherein the vector layer of the first scoring model is the same as the vector layer of the second segmentation model, and the fully connected layer of the first scoring model is used to generate a segmentation score of the corresponding sentence according to the input feature vector.
[0096] like Figures 4 to 5 As shown, in step S3, a first scoring model is constructed based on the second segmentation model, including steps S31 to S33:
[0097] Step S31: extracting the vector layer of the second segmentation model;
[0098] Step S32: construct a fully connected layer, which is used to generate a segmentation score of the corresponding sentence according to the input feature vector;
[0099] Step S33: construct a first scoring model based on the vector layer and the fully connected layer.
[0100] Specifically, when constructing the first scoring model, the vector layer of the second segmentation model is first extracted, and then a fully connected layer is constructed, and the fully connected layer is connected to the vector layer of the second segmentation model, and the vector layer and the fully connected layer constitute the first scoring model.
[0101] It should be noted that the fully connected layer is used to generate the segmentation score of the corresponding sentence based on the input feature vector. Specifically, each node in the website of the fully connected layer is connected to all the nodes of the previous vector layer to integrate the vector features extracted by the vector layer, that is, to map the output of the vector layer to a scalar value, which is the segmentation score of a single sentence in the input sample video.
[0102] Step S4: training the first score model according to the reference segmentation point to generate a second score model.
[0103] See also Figure 6 , Figure 6 A schematic diagram of training a first scoring model in a model training method provided in an embodiment of the present application.
[0104] In some implementations, step S4 trains the first score model according to the reference segmentation point to generate a second score model, including:
[0105] Assign corresponding reference scores to sentences in the sample video according to the reference segmentation points marked in the sample video, where the reference scores represent the distance from the sentences to the reference segmentation points;
[0106] Establishing a second loss function between the reference score and the segmentation score, the second loss function is used to characterize the error of the segmentation score of the sentence relative to the reference score;
[0107] The sample video is input into the first scoring model so that the first scoring model generates a segmentation score, and the parameters of the second scoring model are optimized according to the second loss function until the second loss function converges to obtain the second scoring model.
[0108] It should be noted that the reference score represents the distance from a sentence to the reference segmentation point closest to the sentence. The closer the distance from any sentence in the sample video to the reference segmentation point, the higher the reference score assigned to the sentence.
[0109] It should also be noted that the second loss function is used to characterize the error of the segmentation score relative to the reference score, that is, the difference between the segmentation scores corresponding to multiple sentences in the sample video and the reference score. In the process of training the first scoring model, the output of the second loss function changes accordingly, and the server also optimizes the model parameters of the first score according to the output of the instant second loss function. When the second loss function converges, it indicates that the output of the second loss function is stable, the training of the first scoring model is completed, and the second scoring model is obtained.
[0110] For example, establishing the second loss function may be establishing a pairing ranking loss function based on the time difference between the reference score and the segmentation score in the sample video.
[0111] Specifically, training the first segmentation model to generate the second segmentation model is first to establish a second loss function between the reference score and the segmentation score, then input the sample video into the first scoring model so that the first scoring model generates a segmentation score, and the second loss function is updated according to the generated segmentation score, after which the parameters of the first scoring model are optimized according to the output of the second loss function to iteratively train the first scoring model to obtain the second scoring model.
[0112] After optimizing the parameters of the first scoring model and under the condition that the second loss function has not converged, the server executing the present method will again input the sample video into the first scoring model so that the first scoring model generates a segmentation score, and update the second loss function according to the segmentation score generated by the first scoring model. Thereafter, the parameters of the first scoring model are optimized according to the output of the second loss function to iteratively train the first scoring model until the second loss function converges to obtain the second scoring model.
[0113] It should be noted that in the process of training the first segmentation model to obtain the second segmentation model, the parameters of the vector layer and the segmentation layer in the first segmentation model are adjusted. After the parameters of the vector layer and the segmentation layer are adjusted, the fully connected layer is connected to the vector layer in the second segmentation model to form the first scoring model. In the process of training the first scoring model to obtain the second scoring model, the parameters of the fully connected layer in the first scoring model are adjusted, so that the segmentation score assigned by the second scoring model to each sentence in the sample video is closer to the reference score, thereby improving the accuracy of the second scoring model.
[0114] Step S5: input the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence.
[0115] In some implementations, step S5 inputs the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence, including:
[0116] Inputting the sample video into at least one of a second score model and a second segmentation model to generate a feature vector of the sentence;
[0117] Inputting the sample video into a second scoring model to generate a segmentation score for the sentence;
[0118] Inputting the sample video into a second segmentation model to generate model segmentation points of the sample video;
[0119] The feature vector, segmentation score, and information about whether the sentence is marked with a model segmentation point are integrated to generate a sentence dataset corresponding to the sentence.
[0120] It should be noted that the second scoring model includes a vector layer and a fully connected layer. By inputting the sample video into the second scoring model, the feature vector of the sentence in the sample video and the segmentation score of the sentence generated according to the feature vector can be obtained. The second segmentation model includes a vector layer and a segmentation layer. By inputting the sample video into the second scoring model, the feature vector of the sentence in the sample video and the model segmentation points marked in the sample video according to the feature vectors of multiple sentences can be obtained.
[0121] It should also be noted that the sentence data set corresponding to the sentence includes the sentence's feature vector, segmentation score, and whether the sentence is marked with a model segmentation point. If the sample video is input into the second segmentation model, if the second segmentation model marks a model segmentation point after a certain sentence, it belongs to "the sentence is marked with a model segmentation point"; if the second segmentation model does not mark a model segmentation point after a certain sentence, it belongs to "the sentence is not marked with a model segmentation point".
[0122] Exemplarily, in the process of generating a sentence data set, it is predetermined that “1” represents that a model segmentation point is marked after a sentence, and “0” represents that a model segmentation point is not marked after a sentence.
[0123] Taking a sentence x1 in the sample video as an example, the feature vector of sentence x1 is s1, the segmentation score of sentence x1 is r1, and the model segmentation point is marked after sentence x1, then the sentence data set corresponding to sentence s1 is (s1, 1, r1); taking another sentence x2 in the sample video as an example, the feature vector of sentence x2 is s2, the segmentation score of sentence x2 is r2, and the model segmentation point is marked after sentence x2, then the sentence data set corresponding to sentence s2 is (s2, 0, r2).
[0124] By integrating the feature vectors, segmentation scores, and information about whether a model segmentation point is marked at the sentence of multiple sentences, a sentence dataset corresponding to the sentences in the sample video is generated.
[0125] Step S6: construct an objective function, the output of which matches the expected value of the segmentation score and the update step length of the second segmentation model.
[0126] Step S7: Train the second segmentation model according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
[0127] It should be noted that the process of training the second segmentation model is performed according to the output of the objective function, wherein the training goal of the second segmentation model is to maximize the output of the objective function.
[0128] Specifically, the output of the objective function is positively correlated with the expected value of the segmentation score, and negatively correlated with the update step size of the second segmentation model.
[0129] For example, the objective function is objective=D π [r-βlog(π 1 / π 2 )], where objective represents the objective function, E represents the expectation, D represents the training data set, r represents the expectation of the segmentation score in the sentence data set corresponding to the sentence, and π 1 / π 2 is the update step size of the second segmentation model of the second segmentation model.
[0130] Using this objective function formula to train the second segmentation model, the second segmentation model can learn to assign higher segmentation scores to sentences in the video that are closer to the time point when the content theme changes. The specific training goal of the second segmentation model is to make objective = D π [r-βlog(π 1 / π 2 )] is maximized.
[0131] It should be noted that β is a hyperparameter, which is set manually and generally takes a value of [0,1]. Calculate the value of β and use the gradient ascent function To train the second segmentation model. Wherein, θ is the model parameter in the second segmentation model. In the process of training the second segmentation model, the model parameter θ is iteratively optimized.
[0132] In some implementations, the output of the objective function is proportional to the expected value of the segmentation score and inversely proportional to the update step size of the second segmentation model; then in step S7, the second segmentation model is trained according to the sentence data set until the output of the objective function converges, including:
[0133] The second segmentation model is copied to generate a third segmentation model, and the update step size is the strategy difference between the segmentation points of the model generated by the second segmentation model and the segmentation points of the model generated by the third segmentation model;
[0134] Optimizing the strategy of generating model segmentation points by the third segmentation model according to the sentence data set;
[0135] The model parameters of the third segmentation model are copied to the second segmentation model, so that the second segmentation model adopts the same strategy of generating model segmentation points as the third segmentation model, until the output of the objective function is greater than a preset threshold.
[0136] Specifically, the device executing this method copies the second segmentation model to generate a third segmentation model. It should be noted that the update step size is the difference in strategy between the model segmentation points generated by the second segmentation model and the model segmentation points generated by the third segmentation model. Then, the strategy for generating model segmentation points by the third segmentation model is optimized according to the sentence data set, and then the model parameters of the third segmentation model are copied to the second segmentation model, so that the second segmentation model adopts the same strategy for generating model segmentation points as the third segmentation model, until the output of the objective function converges, and the third segmentation model is trained as the target segmentation model.
[0137] It should be noted that after the second segmentation model is copied to generate the third segmentation model, the strategy of the third segmentation model generating the model segmentation point is updated, while the strategy of the second segmentation model generating the model segmentation point remains unchanged temporarily. Therefore, the strategy of the second segmentation model generating the model segmentation point and the strategy of the third segmentation model generating the model segmentation point differ. That is, in the aforementioned objective function formula objective=D π [r-βlog(π 1 / π 2 )], π 1 Generate the probability distribution of the model segmentation points for the second segmentation model, π 2 Generates the probability distribution of the model segmentation points for the third segmentation model, and π 1 / π 2 It is the ratio of the probability distribution and also represents the update step size of the second segmentation model.
[0138] In summary, the present application obtains a first segmentation model, a training data set including a sample video and a reference segmentation point; trains the first segmentation model according to the reference segmentation point to generate a second segmentation model; constructs a first scoring model for assigning segmentation scores to sentences in the sample video based on the second segmentation model; trains the first scoring model according to the reference segmentation point to generate a second scoring model; generates a sentence data set according to the training data set; constructs an objective function, whose output matches the expected value of the segmentation score and the update step size of the second segmentation model; trains the second segmentation model according to the sentence data set until the output of the objective function converges, and obtains a target segmentation model that can segment the video to be edited input into the model based on the content theme of the sample video, thereby improving the model accuracy, the accuracy of video content theme recognition, and the film-making effect of video editing, so that the cutting points output by the model when processing the video to be edited are closer to the segmentation points actually required by the user.
[0139] See also Figure 7 , Figure 7 A flowchart of a video generation method provided in an embodiment of the present application.
[0140] like Figure 7 As shown, the embodiment of the present application also provides a video generation method, the method comprising steps S81-S83:
[0141] Step S81: receiving a video to be edited and an editing instruction;
[0142] Step S82: calling a target segmentation model according to the clipping instruction, wherein the target segmentation model is trained by using any model training method provided in the embodiments of the present application;
[0143] Step S83: input the video to be edited into the target segmentation model so that the target segmentation model generates model segmentation points corresponding to the video to be edited, and segments the video to be edited according to the model segmentation points to generate at least one target video.
[0144] Specifically, the video generation method provided in the embodiment of the present application first receives the video to be edited and the editing instructions, and then calls the target segmentation model according to the editing instructions, wherein the target segmentation model is trained using any one of the model training methods provided in the embodiment of the present application, and the video to be edited is input into the target segmentation model, so that the target segmentation model generates model segmentation points corresponding to the video to be edited, so as to segment the video to be edited according to the model segmentation points, thereby generating at least one target video.
[0145] Since the model training method provided in the embodiment of the present application is to train the first segmentation model multiple times based on the training data set, the obtained target segmentation model can segment the video to be edited input into the model based on the content theme of the sample video, thereby improving the model accuracy, the accuracy of video content theme recognition and the film effect of video editing, so that the cutting points output by the model when processing the video to be edited are closer to the segmentation points actually required by the user.
[0146] See also Figure 8 , Figure 8 A schematic block diagram of the structure of the model training device provided in an embodiment of the present application.
[0147] like Figure 8 As shown, the model training device 100 can be applied to a computer device, and the model training device 100 includes a task division module 101, a record acquisition module 102, a task division module 103, a task screening module 104, and a task allocation module 105.
[0148] See also Fig. 9 , Fig. 9 A schematic diagram of the module structure of a model training device provided in an embodiment of the present application.
[0149] like Fig. 9 As shown, the embodiment of the present application also provides a model training device 100, including:
[0150] The sample acquisition module 101 is used to acquire a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points;
[0151] A first training module 102, configured to train the first segmentation model according to the reference segmentation point to generate a second segmentation model;
[0152] A scoring model module 103, configured to construct a first scoring model based on the second segmentation model, wherein the first scoring model is configured to assign segmentation scores to sentences in the sample video, wherein the segmentation scores represent the distance between the sentences and the model segmentation points;
[0153] A second training module 104, configured to train the first scoring model according to the reference segmentation point to generate a second scoring model;
[0154] The sample processing module 105 is used to input the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence;
[0155] A function construction module 106 is used to construct an objective function, the output of which matches the expected value of the segmentation score and the learning rate of the second segmentation model;
[0156] The third training module 107 is used to train the second segmentation model according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
[0157] In some embodiments, the output of the objective function is proportional to the expected value of the segmentation score and inversely proportional to the update step size of the second segmentation model;
[0158] The third training module 107 trains the second segmentation model according to the sentence data set until the output of the objective function converges, including:
[0159] The second segmentation model is copied to generate a third segmentation model, and the update step size is the strategy difference between the segmentation points of the model generated by the second segmentation model and the segmentation points of the model generated by the third segmentation model;
[0160] Optimizing the strategy of generating model segmentation points by the third segmentation model according to the sentence data set;
[0161] The model parameters of the third segmentation model are copied to the second segmentation model, so that the second segmentation model adopts the same strategy of generating model segmentation points as the third segmentation model, until the output of the objective function converges.
[0162] In some embodiments, when the first training module 102 trains the first segmentation model according to the reference segmentation point to generate the second segmentation model, the first training module 102 includes:
[0163] Establishing a first loss function between the reference segmentation point and the model segmentation point, the first loss function is used to characterize the error of the model segmentation point relative to the reference segmentation point;
[0164] The sample video is input into the first segmentation model so that the first segmentation model generates model segmentation points, and the parameters of the first segmentation model are optimized according to the first loss function until the first loss function converges, thereby obtaining a second segmentation model.
[0165] In some embodiments, the second segmentation model includes a vector layer and a segmentation layer, the vector layer is used to obtain a feature vector of a sentence in a sample video, and the segmentation layer is used to generate a model segmentation point according to the feature vector of the sentence;
[0166] When constructing the first scoring model based on the second segmentation model, the scoring model module 103 includes:
[0167] Building a first scoring model based on the second segmentation model includes:
[0168] Extracting a vector layer of a second segmentation model;
[0169] Construct a fully connected layer, which is used to generate the segmentation score of the corresponding sentence based on the input feature vector;
[0170] The first scoring model is constructed based on the vector layer and the fully connected layer.
[0171] In some implementations, when the second training module 104 trains the first score model according to the reference segmentation point to generate the second score model, the second training module 104 includes:
[0172] Assign corresponding reference scores to sentences in the sample video according to the reference segmentation points marked in the sample video, where the reference scores represent the distance from the sentences to the reference segmentation points;
[0173] Establishing a second loss function between the reference score and the segmentation score, the second loss function is used to characterize the error of the segmentation score of the sentence relative to the reference score;
[0174] The sample video is input into the first scoring model so that the first scoring model generates a segmentation score, and the parameters of the second scoring model are optimized according to the second loss function until the second loss function converges to obtain the second scoring model.
[0175] In some embodiments, when the sample processing module 105 inputs the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to the sentences one by one, it includes:
[0176] Inputting the sample video into at least one of a second score model and a second segmentation model to generate a feature vector of the sentence;
[0177] Inputting the sample video into a second scoring model to generate a segmentation score for the sentence;
[0178] Inputting the sample video into a second segmentation model to generate model segmentation points of the sample video;
[0179] The feature vector, segmentation score, and information about whether the sentence is marked with a model segmentation point are integrated to generate a sentence dataset corresponding to the sentence.
[0180] See also Fig. 9 , Fig. 9 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application.
[0181] like Fig. 9 As shown, an embodiment of the present application further provides a computer device 200, which includes a processor 201 and a memory 202, and the processor 201 and the memory 202 are connected via a bus 203, such as an I2C (Inter-integrated Circuit) bus.
[0182] Specifically, the processor 201 is used to provide computing and control capabilities to support the operation of the entire computer device. The processor 301 can be a central processing unit (CPU), and the processor 301 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0183] Specifically, the memory 202 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.
[0184] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is merely a block diagram of a partial structure related to the embodiment of the present application, and does not constitute a limitation on the computer device to which the embodiment of the present application is applied. Specifically, the computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0185] Among them, the processor 301 is used to run the computer program stored in the memory, and implement any model training method provided in the embodiments of the present application when executing the computer program.
[0186] In some implementations, the processor 301 is configured to run a computer program stored in the memory, and implement the following steps when executing the computer program:
[0187] Obtaining a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points;
[0188] Training the first segmentation model according to the reference segmentation point to generate a second segmentation model;
[0189] Building a first scoring model based on the second segmentation model, the first scoring model is used to assign segmentation scores to sentences in the sample video, the segmentation scores representing the distance from the sentence to the model segmentation point;
[0190] Training the first scoring model according to the reference segmentation point to generate a second scoring model;
[0191] Inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence;
[0192] Constructing an objective function, the output of which matches the expected value of the segmentation score and the update step size of the second segmentation model;
[0193] The second segmentation model is trained according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
[0194] In some implementations, the output of the objective function is proportional to the expected value of the segmentation score and inversely proportional to the update step size of the second segmentation model. When the processor 301 trains the second segmentation model according to the sentence data set until the output of the objective function converges, the processor 301 includes:
[0195] The second segmentation model is copied to generate a third segmentation model, and the update step size is the strategy difference between the segmentation points of the model generated by the second segmentation model and the segmentation points of the model generated by the third segmentation model;
[0196] Optimizing the strategy of generating model segmentation points by the third segmentation model according to the sentence data set;
[0197] The model parameters of the third segmentation model are copied to the second segmentation model, so that the second segmentation model adopts the same strategy of generating model segmentation points as the third segmentation model, until the output of the objective function converges.
[0198] In some implementations, when the processor 301 trains the first segmentation model according to the reference segmentation point to generate the second segmentation model, the processor 301 includes:
[0199] Establishing a first loss function between the reference segmentation point and the model segmentation point, the first loss function is used to characterize the error of the model segmentation point relative to the reference segmentation point;
[0200] The sample video is input into the first segmentation model so that the first segmentation model generates model segmentation points, and the parameters of the first segmentation model are optimized according to the first loss function until the first loss function converges, thereby obtaining a second segmentation model.
[0201] In some embodiments, the second segmentation model includes a vector layer and a segmentation layer, the vector layer is used to obtain a feature vector of a sentence in a sample video, and the processor 301 includes:
[0202] Building a first scoring model based on the second segmentation model includes:
[0203] Extracting a vector layer of a second segmentation model;
[0204] Construct a fully connected layer, which is used to generate the segmentation score of the corresponding sentence based on the input feature vector;
[0205] The first scoring model is constructed based on the vector layer and the fully connected layer.
[0206] In some implementations, when the processor 301 trains the first score model according to the reference segmentation point to generate the second score model, the processor 301 includes:
[0207] Assign corresponding reference scores to sentences in the sample video according to the reference segmentation points marked in the sample video, where the reference scores represent the distance from the sentences to the reference segmentation points;
[0208] Establishing a second loss function between the reference score and the segmentation score, the second loss function is used to characterize the error of the segmentation score of the sentence relative to the reference score;
[0209] The sample video is input into the first scoring model so that the first scoring model generates a segmentation score, and the parameters of the second scoring model are optimized according to the second loss function until the second loss function converges to obtain the second scoring model.
[0210] In some implementations, when the processor 301 inputs the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to each sentence, the processor 301 includes:
[0211] Inputting the sample video into at least one of a second score model and a second segmentation model to generate a feature vector of the sentence;
[0212] Inputting the sample video into a second scoring model to generate a segmentation score for the sentence;
[0213] Inputting the sample video into a second segmentation model to generate model segmentation points of the sample video;
[0214] The feature vector, segmentation score, and information about whether the sentence is marked with a model segmentation point are integrated to generate a sentence dataset corresponding to the sentence.
[0215] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the computer device described above can refer to the corresponding process in the aforementioned model training method embodiment, and will not be repeated here.
[0216] An embodiment of the present application also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any model training method provided in the embodiments of the present application specification.
[0217] The storage medium may be an internal storage unit of the computer device of the aforementioned embodiment, such as a hard disk or memory of the computer device. The storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device.
[0218] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0219] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0220] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A model training method, characterized in that: The method comprises: Acquire a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points; Training the first segmentation model according to the reference segmentation point to generate a second segmentation model; Building a first scoring model based on the second segmentation model, the first scoring model is used to assign a segmentation score to the sentence in the sample video, the segmentation score representing the distance from the sentence to the segmentation point of the model; Training the first score model according to the reference segmentation point to generate a second score model; Inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to the sentences one by one; Constructing an objective function, wherein an output of the objective function matches an expected value of the segmentation score and an update step length of the second segmentation model; The second segmentation model is trained according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
2. The method according to claim 1, characterized in that The output of the objective function is proportional to the expected value of the segmentation score and inversely proportional to the update step size of the second segmentation model, and the second segmentation model is trained according to the sentence data set until the output of the objective function converges, including: The second segmentation model is copied to generate a third segmentation model, and the update step length is the difference between the strategies of generating the model segmentation points by the second segmentation model and generating the model segmentation points by the third segmentation model; Optimizing the strategy of generating the model segmentation points by the third segmentation model according to the sentence data set; The model parameters of the third segmentation model are copied to the second segmentation model, so that the second segmentation model adopts the same strategy of generating the model segmentation points as the third segmentation model until the output of the objective function converges.
3. The method according to claim 1, characterized in that The training of the first segmentation model according to the reference segmentation point to generate a second segmentation model includes: Establishing a first loss function between the reference segmentation point and the model segmentation point, wherein the first loss function is used to characterize an error of the model segmentation point relative to the reference segmentation point; The sample video is input into the first segmentation model so that the first segmentation model generates the model segmentation points, and the parameters of the first segmentation model are optimized according to the first loss function until the first loss function converges, thereby obtaining the second segmentation model.
4. The method according to claim 1, characterized in that: The second segmentation model includes a vector layer and a segmentation layer, the vector layer is used to obtain the feature vector of the sentence in the sample video, and the segmentation layer is used to generate the model segmentation point according to the feature vector of the sentence; The constructing a first scoring model based on the second segmentation model comprises: extracting the vector layer of the second segmentation model; Constructing a fully connected layer, wherein the fully connected layer is used to generate the segmentation score corresponding to the sentence according to the input feature vector; The first scoring model is constructed according to the vector layer and the fully connected layer.
5. The method according to claim 4, characterized in that The training the first score model according to the reference segmentation point to generate a second score model includes: Assigning a corresponding reference score to the sentence in the sample video according to a reference segmentation point marked in the sample video, wherein the reference score represents a distance from the sentence to the reference segmentation point; Establishing a second loss function between the reference score and the segmentation score, wherein the second loss function is used to characterize an error of the segmentation score of the sentence relative to the reference score; The sample video is input into the first score model so that the first score model generates the segmentation score, and the parameters of the second score model are optimized according to the second loss function until the second loss function converges to obtain the second score model.
6. The method according to claim 1, characterized in that Inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to the sentences one by one, including: Inputting the sample video into at least one of the second score model and the second segmentation model to generate a feature vector of the sentence; Inputting the sample video into the second scoring model to generate a segmentation score for the sentence; Inputting the sample video into the second segmentation model to generate model segmentation points of the sample video; The feature vector of the sentence, the segmentation score, and information on whether the model segmentation point is marked at the sentence are integrated to generate a sentence data set corresponding to the sentence.
7. A video generation method, characterized in that: The method comprises: Receive the video to be edited and the editing instruction; Calling a target segmentation model according to the clipping instruction, wherein the target segmentation model is trained using the model training method according to any one of claims 1 to 6; The video to be edited is input into the target segmentation model so that the target segmentation model generates model segmentation points corresponding to the video to be edited, and the video to be edited is segmented according to the model segmentation points to generate at least one target video.
8. A model training device, characterized in that: include: A sample acquisition module, used to acquire a first segmentation model and a training data set, wherein the training data set includes a sample video and reference segmentation points annotated in the sample video, and the first segmentation model is used to generate model segmentation points to segment the sample video according to the model segmentation points; A first training module, configured to train the first segmentation model according to the reference segmentation point to generate a second segmentation model; A scoring model module, configured to construct a first scoring model based on the second segmentation model, wherein the first scoring model is configured to assign a segmentation score to a sentence in the sample video, wherein the segmentation score represents a distance from the sentence to a segmentation point of the model; A second training module, configured to train the first score model according to the reference segmentation point to generate a second score model; A sample processing module, used for inputting the training data set into the second scoring model and the second segmentation model respectively to generate a sentence data set corresponding to the sentences one by one; A function construction module, used for constructing an objective function, the output of which matches the expected value of the segmentation score and the learning rate of the second segmentation model; The third training module is used to train the second segmentation model according to the sentence data set until the output of the objective function converges to obtain a target segmentation model.
9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the model training method as described in any one of claims 1 to 6 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the model training method as described in any one of claims 1 to 6.