Method, device, storage medium and program product for processing line text

By identifying and processing long text statements in video line text, and using the large language model to generate reference text statements for division, the problem of "AI illusion" that may occur in line text generation by large language model is solved, and the efficiency and accuracy of video production are improved.

CN118714416BActive Publication Date: 2025-05-30ALI HEALTH TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411197703.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-05-30
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

In the process of automatically generating lines and texts using large language models, there may be an 'AI hallucination' problem, which will affect the efficiency of video production.

Method used

By obtaining the line text corresponding to the video, identifying and processing long text statements, calling the large language model to generate reference text statements, and dividing long text statements into standard text statements suitable for the video according to these reference text statements.

Benefits of technology

It effectively overcomes the problem of ‘AI illusion’, reduces the workload of errors and corrections in video production, and improves the efficiency and accuracy of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118714416B_ABST
    Figure CN118714416B_ABST
Patent Text Reader

Abstract

An embodiment of this specification provides a method, device, storage medium, and program product for processing a line text. The method includes: obtaining the line text corresponding to a video; wherein the line text includes a plurality of text sentences; the text sentences include long text sentences and standard text sentences; the number of words in the standard text sentence is within a specified word count threshold range; the number of words in the long text sentence exceeds the specified word count threshold range; calling a large language model based on the long text sentence to obtain a plurality of reference text sentences corresponding to the long text sentence output by the large language model; dividing the long text sentence into standard text sentences applicable to the video according to the text of the long text sentence included in the plurality of reference text sentences; wherein the standard text sentence is used as the subtitle of the video. Through the embodiment of this specification, on the basis of using the large language model to improve the video production efficiency, the influence of the "AI hallucination" problem on the video production efficiency can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the technical field of video processing, and particularly to a method, device, storage medium, and program product for processing line text. Background Art

[0002] In the digital media era, using videos to spread medical and health knowledge has become an important form of knowledge popularization in the medical field, which has a positive impact on the dissemination efficiency of medical information and the acceptance of patients by patients. The effective conveyance of video content depends on visual elements (such as line text) and auditory elements (such as video voice). Due to the strong professionalism of medical and health knowledge, the accuracy and readability of line text play an important role in the dissemination process of popular science videos.

[0003] To reduce the difficulty for practitioners in the medical and health field to produce medical popular science videos and alleviate the workload of producing medical popular science videos, large language model (LLM) technology has been widely applied. Specifically, traditional methods for processing line text usually rely on manual editing, which is time-consuming and laborious, and prone to errors when dealing with information in professional fields, such as medical terms. With the iteration of artificial intelligence technology, it is possible to use large language models to achieve the automatic generation of line text.

[0004] The key to automatic subtitle generation lies in the division of long texts. However, when using large language models to process long text sentences, there may be an "AI hallucination" problem. In professional fields such as medical and health, the "AI hallucination" problem may seriously affect the production efficiency of videos. Summary of the Invention

[0005] In view of this, multiple embodiments of this specification are committed to providing a method, device, storage medium, and program product for processing line text, which can overcome the "AI hallucination" problem in the process of using large language models to achieve the automatic generation of line text.

[0006] An embodiment of this specification provides a method for processing line text, including: obtaining the line text corresponding to a video; wherein, the line text includes multiple text sentences; the text sentences include long text sentences and standard text sentences; wherein, the number of words in the standard text sentences is within a specified word count threshold range; the number of words in the long text sentences exceeds the specified word count threshold range; calling a large language model based on the long text sentences to obtain multiple reference text sentences corresponding to the long text sentences output by the large language model; dividing the long text sentences into standard text sentences applicable to the video according to the words of the long text sentences included in the multiple reference text sentences; wherein, the standard text sentences serve as the subtitle text of the video.

[0007] One embodiment of this specification provides a video production system, including: an acquisition module, configured to acquire the line text corresponding to a video; wherein, the line text includes a plurality of text sentences; the text sentences include long text sentences and standard text sentences; wherein, the number of words in the standard text sentence is within a specified word count threshold range; the number of words in the long text sentence exceeds the specified word count threshold range; a call module, configured to call a large language model based on the long text sentence to obtain a plurality of reference text sentences corresponding to the long text sentence output by the large language model; a division module, configured to divide the long text sentence into standard text sentences applicable to the video according to the text of the long text sentence included in the plurality of reference text sentences; wherein, the standard text sentence serves as the subtitle text of the video.

[0008] One embodiment of this specification provides a computer device, which includes a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the processing method of the line text described in any of the above embodiments.

[0009] One embodiment of this specification provides a computer-readable storage medium, in which at least one computer program is stored. When the at least one computer program is executed by a processor, it can implement the processing method of the line text described in any of the above embodiments.

[0010] One embodiment of this specification provides a computer program product, which is used to implement the processing method of the line text described in any of the above embodiments.

[0011] Multiple embodiments provided in this specification obtain the long text sentences included in the line text corresponding to a video, call a large language model based on the long text sentence, obtain a plurality of reference text sentences corresponding to the long text sentence output by the large language model, and then divide the long text sentence into standard text sentences applicable to the video according to the text of the long text sentence included in the plurality of reference text sentences, so as to generate the line text according to the standard text sentences. In this way, during the process of video production, the user can use the large language model to determine the division position of the long text sentence in the video line text, and then divide the long text sentence according to the division position, so as to reduce the impact of the "AI hallucination" problem on the video production efficiency. Description of the Drawings

[0012] Figure 1 Flowchart of using digital humans to generate videos provided for related technologies.

[0013] Figure 2Schematic diagram of generating a video using the method for processing line text provided in one embodiment of this specification.

[0014] Figure 3 Flow schematic diagram of the method for processing line text provided in one embodiment of this specification.

[0015] Figure 4 Schematic diagram of the video production system provided in one embodiment of this specification.

[0016] Figure 5 Schematic diagram of the computer device provided in one embodiment of this specification. Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0018] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0019] In fields with strong professionalism such as medical and health, popular science videos have now become an important way to disseminate professional knowledge to the public. Taking the medical and health field as an example, the popular science videos in this field are mainly produced and released by the practitioners in this field. For example, doctors, nurses, pharmacists, pharmaceutical R & D personnel, etc. However, due to the limitations of their professional characteristics, the time that practitioners in this field can spend on the production of popular science videos is relatively limited. Therefore, digital human technology has been widely used in the production process of popular science videos in this field. By creating a simulation "digital human" corresponding to the video producer, the digital human can be used to replace the video producer for voice broadcasting of video content. That is, the video producer only needs to upload the line text corresponding to the video to the video production system, and the video production system can automatically generate popular science videos, thereby reducing the difficulty and workload of video production by the video producer and improving the production efficiency of popular science videos.

[0020] Please refer to Figure 1。Currently, the process of making popular science videos using digital human technology mainly includes: video line text submission, video line text correction and subtitle text generation, digital human video generation, video cover frame production, line text and picture matching, and video review and release. In the videos made using digital human technology, the line text is usually located below the digital human image, that is, in the bottom area of the video screen. Due to the size limitation of the bottom area of the video screen, during the line text generation process, it is necessary to divide long text sentences in the video line text that exceed the specified word count threshold range.

[0021] In related technologies, the division of long text sentences in video line text is mainly completed by relying on the subtitle automatic generation plug-in integrated in the video production system or an independent subtitle automatic generation software and the manual correction of video producers. Specifically, video producers can submit the video line text to the subtitle automatic generation plug-in or subtitle automatic generation software. The subtitle automatic generation plug-in or subtitle automatic generation software can combine the long text division rules set by video producers and common natural language processing text division rules to divide the long text sentences in the video line text and output the division results. After obtaining the long text sentence division results output by the subtitle automatic generation plug-in or subtitle automatic generation software, video producers can perform manual correction on the long text sentence division results and use the corrected long text sentence division results as the line text. The correction content can include: splicing terms that are mis-divided into multiple short texts by the subtitle automatic generation software, adjusting short texts with inconsistent semantics in the long text division results, etc.

[0022] For fields with strong professionalism such as medical and health, since there may be many rare professional terms involved, relying solely on the voice broadcast of digital humans, it may be difficult for viewers to understand the content of popular science videos or there may be ambiguity in the content of popular science videos. At this time, viewers usually combine the voice broadcast of digital humans with video subtitles to accurately understand the video content. Therefore, the knowledge accuracy of video subtitles plays an important role in achieving the overall knowledge accuracy of popular science videos. However, due to the limitation of the development of natural language processing technology, it is difficult to accurately divide long text sentences in video line text only by relying on subtitle automatic generation plug-ins or subtitle automatic generation software. Video producers need to spend a lot of time and effort correcting the long text sentence division results output by subtitle automatic generation plug-ins or subtitle automatic generation software, resulting in an increase in the cost and a decrease in the efficiency of video production.

[0023] The emergence of large language models has helped improve the production efficiency of popular science videos in the medical and health fields. Compared with the subtitle automatic generation plug-ins or software in related technologies, large language models have stronger semantic understanding and natural language processing capabilities. Therefore, video producers can use large language models to replace subtitle automatic generation plug-ins or software to divide long text sentences in video line text, that is, input long text sentences in video line text into large language models, and use the output of large language models as the division result of long text sentences, thereby reducing the time spent by video producers on correcting the division results of long text sentences and shortening the overall time required for video producers to produce videos.

[0024] With the wide application of large language models, video producers have found that in the division results of long text sentences output by large language models, there may be fabrications and tampering of long text sentences, which can be called the "AI hallucination" problem. For example, inputting the long text sentence "Gene detection, this super detective appears" into a large language model, the output of the large language model may be "Gene detection", "This detective leaves", or "Gene detection", "This police officer appears". If the output result of the large language model is directly used as the video subtitle, it may mislead the audience, and for popular science videos in the medical and health fields, it may even cause group health risks. To reduce the impact of the "AI hallucination" problem on the knowledge accuracy of popular science videos, video producers may still need to spend a lot of time and effort to correct the division results of long text sentences output by large language models, without giving full play to the role of large language models in improving video production efficiency.

[0025] Therefore, it is necessary to provide a method for processing line text to overcome the "AI hallucination" problem on the basis of using large language models to divide long text sentences and improve video production efficiency.

[0026] Please refer to Figure 2 . This specification provides an application scenario example of a method for processing line text. This method for processing line text can be applied to a video production system that uses digital human technology to produce videos. After generating a digital human image corresponding to the user, the user only needs to submit the video line text to this video production system, and this video production system can automatically generate a video in which the digital human broadcasts the content of the line text.

[0027] Taking the production of popular science videos in the medical and health fields using the above video production system as an example, on the basis of having generated a digital human image corresponding to the user, the process of automatically generating popular science videos using this video production system is as follows.

[0028] The user can submit a draft line text corresponding to the video to the video production system. This draft line text can include multiple long text sentences.

[0029] Since the digital human will perform voice announcements based on the content of the line text, to improve the presentation effect of the line text and to ensure the consistency between the voice announcement content of the digital human and the line text content, after the video production system receives the draft line text, the draft line text is stored in the database for reading and calling during subsequent processing.

[0030] The video production system can obtain the draft line text from the database and perform preprocessing on the draft line text. Specifically, the content of the preprocessing of the line text specifically includes processing such as number unification, bracket removal, pause time setting, and character replacement, and outputs the line text. Among them, the number unification processing can unify the numbers in various formats such as Roman numerals, Chinese numerals, Arabic numerals, and English letters in the draft line text into Arabic numeral format numbers. The bracket removal processing can remove the brackets and the text used for supplementation or explanation enclosed by the brackets in the draft line text. The pause time setting processing can set the time interval for the digital human to announce different sentences according to the punctuation marks between different sentences in the draft line text. The character replacement processing can replace punctuation marks such as double quotes in the draft line text with punctuation marks such as commas. The content of the processed line text is consistent with the content of the voice announcement of the digital human in the automatically generated video. After the preprocessing of the line text, the voice announcement effect of the digital human will be closer to the effect of a real person's narration.

[0031] The video production system can further perform text division processing on the long text sentences in the line text to finally obtain text sentences suitable for the video. The text sentences in the line text can include the theme text and the subtitle text. The theme text can be used as the theme of the video, and the subtitle text is used as the video subtitle of the video.

[0032] Take the division processing of the subtitle text as an example. For example, the subtitle text is "When a patient with viral cold strays into the neurology consulting room". At this time, the subtitle text has exceeded the word count threshold of 10 words that a single-line title can accommodate. Then, a prompt instruction including this subtitle text can be constructed to instruct the large language model to output the large model line-breaking result corresponding to this subtitle text. The large model line-breaking result can include multiple reference text sentences. The multiple reference text sentences included in the large model line-breaking result can be "When a patient with viral rhinitis catches a cold" and "Strays into the neurology and neurosurgery consulting room" respectively. Comparing the subtitle text with the large model line-breaking result, it can be seen that the content of the subtitle text has been modified in the large model line-breaking result. If the large model line-breaking result is directly used as the video subtitle, it may mislead the video audience.

[0033] Therefore, after obtaining the line-breaking results of the large model, it is necessary to use the line-breaking results of the large model to determine the division positions for the subtitle text. Specifically, according to the order of the characters in the theme text, determine the consecutive common text between the reference text sentences and the subtitle text in each line-breaking result of the large model. In this way, according to the sentence division of multiple reference text sentences and the corresponding consecutive common text, determine the division positions of the subtitle text. For example, determine the consecutive common text between the reference text sentences of the line-breaking results of the large model and the subtitle text respectively. The consecutive common text between the reference text sentence "When a patient has viral rhinitis and a cold" and the subtitle text may include "When a viral" and "and a cold patient", and the consecutive common text between the reference text sentence "enters the neurology and neurosurgery consultation room by mistake" and the subtitle text may include "enters the neurology" and "and the consultation room". In this way, "When a viral" and "and a cold patient" correspond to one reference text sentence, and "enters the neurology" and "and the consultation room" correspond to one reference text sentence, so that in order, a division position for the subtitle text is formed between "and a cold patient" and "enters the neurology", so that the subtitle text "When a patient has viral rhinitis and a cold and enters the neurology consultation room by mistake" can be divided into two text sentences: "When a patient has viral rhinitis and a cold" and "enters the neurology consultation room by mistake".

[0034] In some cases, the line-breaking results of the large model may not be very accurate. For example, multiple characters that are a single word may be split into two standard text sentences among multiple reference text sentences. For example, the multiple reference text sentences included in the line-breaking results of the large model for the subtitle text are: "When a patient has viral seasonal cold" and "enters the neurology consultation room by mistake". At this time, the consecutive common text between the reference text sentence "When a patient has viral seasonal cold" and the subtitle text is "When a viral" and "seasonal cold", and the consecutive common text between the reference text sentence "enters the neurology consultation room by mistake" and the subtitle text is "enters the neurology" and "and the consultation room". From this, it can be determined that a division position is formed between "seasonal cold" and "enters the neurology", and the subtitle text is divided into two text sentences: "When a patient has viral cold" and "enters the neurology consultation room by mistake". At this time, the word "cold patient" is split and placed in different text sentences respectively, and this division of the subtitle text is unreasonable. At this time, the large language model can be called again using the prompt instruction to regenerate the division results of the large model and re-determine the division positions, etc., to divide the subtitle text into more reasonable multiple text sentences. In some cases, the large language model can be called multiple times to obtain multiple division positions for the subtitle text, and the video production system can determine the most suitable division position among the multiple division positions, so as to obtain the text sentences suitable for the division of the subtitle text.

[0035] The video production system can also perform the foregoing text division process on the theme text to obtain standard text sentences suitable for use as video themes. Specifically, the content of the text division process performed by the video production system on the theme text is similar to the processing of subtitle text, which will not be elaborated here. The video production system can display multiple text sentences after dividing the theme text on the video cover frame.

[0036] When the video production system recognizes that there are short text sentences with fewer characters in the divided text sentences, it can connect the short text sentences with the sentences adjacent to the short text sentences to obtain connected text sentences, thereby reducing the difference in the number of characters between lines of the dialogue text. The video production system can replace special characters that the digital human cannot perform voice announcements on, such as *, #, etc., in the connected text sentences with characters that the digital human can perform voice announcements on. Finally, standard text sentences suitable for the video are obtained. The video production system can use the standard text sentences to generate a video, use the title text of the dialogue text as the title of the video, and use the subtitle text as the video subtitles of the video.

[0037] In summary, the user only needs to provide the draft dialogue text corresponding to the video to the above video production system to obtain a video with the digital human image voice-announcing the dialogue text. Moreover, in this video, the role of the large language model is only to determine the division positions of long text sentences and will not affect the content of the long text sentences, thus overcoming the "AI hallucination" problem in the process of using the large language model to generate dialogue text, reducing the workload of the user in making the video, shortening the video production time, and improving the video production efficiency.

[0038] An embodiment of this specification provides a method for processing dialogue text, and the method for processing dialogue text can be applied to a video production system. The video production system can include a client and a server.

[0039] The client can be an electronic device with network access capabilities. Specifically, for example, the client can be a desktop computer, a tablet computer, a laptop computer, a smart phone, a digital assistant, a smart wearable device, a shopping guide terminal, a television, a smart speaker, a microphone, etc. Among them, smart wearable devices include but are not limited to smart bracelets, smart watches, smart glasses, smart helmets, smart necklaces, etc. Alternatively, the client can also be software that can run on an electronic device.

[0040] The server can be an electronic device with a certain computing and processing capacity. It can have a network communication module, a processor, a memory, etc. The server can also be a distributed server, which can be a system with multiple processors, memories, network communication modules, etc. operating in coordination. Or, the server can also be a server cluster formed by several servers. Or, with the development of science and technology, the server can also be a new technical means capable of realizing the corresponding functions of the embodiments of the specification. For example, it can be a new form of "server" based on quantum computing.

[0041] Please refer to Figure 3 . In this embodiment, the method for processing the line text may include the following steps.

[0042] Step S110: Obtain the line text corresponding to the video.

[0043] In the case where a digital human image corresponding to the user has been generated, for fields with strong professionalism such as medical and health, since the line text is an important way for video viewers to obtain accurate information, improving the accuracy of the line text is very important in the process of generating the line text. To reduce errors in the line text, the video production system can first obtain the line text corresponding to the video that has been proofread.

[0044] In this embodiment, the line text can be a reference text file for the digital human to perform voice broadcast. Specifically, the line text can include multiple text sentences. The text sentences can be formed according to the text content and people's language habits, and different text sentences can be separated by punctuation marks. The line text can include subtitle text used as the subtitle of the video.

[0045] In the video frame, the subtitle text is usually located in a specified area at the bottom of the frame, and the area of the specified area limits the number of words that can be displayed in the specified area. At the same time, to improve the dissemination effect of video information, the number of words included in the subtitle text of each frame of the video should not be too many. For text sentences with a large number of words, it is necessary to divide them for line-by-line display. Therefore, in this embodiment, the text sentences can include long text sentences and standard text sentences. Among them, the number of words in the standard text sentence can be within the specified word count threshold range, while the number of words in the long text sentence exceeds the specified word count threshold range. For example, the specified word count threshold range can include a maximum word count threshold. The number of words in the standard text sentence is less than the maximum word count threshold. The maximum word count threshold can be 9 words, 10 words, 12 words, 13 words, 14 words, 15 words, etc., which will not be elaborated here. The specified word count threshold range can include a minimum word count threshold and a maximum word count threshold. Among them, the minimum word count threshold is less than the maximum word count threshold. The number of words in the standard text sentence is not less than the minimum word count threshold and not greater than the maximum word count threshold.

[0046] Since the digital human can only perform voice broadcasts according to the content of the line text, it is difficult to determine which content in the line text is the main text content to be broadcast and which content is supplementary or explanatory content that does not need to be broadcast. In some cases, directly using the draft line text uploaded by the user to the video production system as the line text for the digital human's voice broadcast may result in a large gap between the digital human's voice broadcast effect and the real person's storytelling effect, affecting the video dissemination effect and the audience's viewing experience. Therefore, the draft line text uploaded by the user to the video production system can be preprocessed.

[0047] In some embodiments, the video production system can identify the note text in the draft line text; remove the note text from the draft line text to obtain the line text.

[0048] In this embodiment, the draft line text can be a text file created by the user and provided to the video production system for the digital human's language broadcast. Specifically, to improve the knowledge accuracy of the line text, the draft line text can be a text file that has been reviewed and proofread and no longer contains knowledge errors.

[0049] In this embodiment, the note text can be text used to express the note content. Specifically, the note text can be text used to supplement and explain the main text. For example, for the text "Common antipyretic analgesics (paracetamol, ibuprofen, and aspirin)" in the draft line text, "Common antipyretic analgesics" is the main text, while "(paracetamol, ibuprofen, and aspirin)" is the note text. The note text can be distinguished from the main text by punctuation marks or special identifiers. For example, the main text and the note text can be distinguished by parentheses, and the text enclosed in parentheses is the note text, and the text not enclosed in parentheses is the main text. It can also be distinguished by the identifier "Note:", and the text following this identifier is the note text.

[0050] In this embodiment, the acquisition of the draft line text can be triggered by the user. Specifically, the user can perform a specified operation on a specified control in the client interface of the video production system. For example, the user can click on the "Upload" control in the client interface of the video production system to upload the draft line text corresponding to the video to the database of the video production system. After the server of the video production system receives the draft line text uploaded by the user, it executes the processing process for the line text. In some embodiments, the video production system can also run entirely on an electronic device that has good data processing capabilities and can be deployed with a large language model. Of course, the large language model can also be deployed on the network side, and this electronic device can call this large language model.

[0051] In this embodiment, the video production system can use punctuation marks or special identifiers that distinguish body text from note text to identify note text in the draft script text. Specifically, after the video production system obtains the draft script text provided by the user, it can traverse the characters in the draft script text. When a specified punctuation mark or a specified identifier is recognized, the text connected or consecutive to the specified punctuation mark or the specified identifier is regarded as note text. For example, when it is recognized that the draft script text contains left and right brackets, the text between the left and right brackets can be regarded as note text.

[0052] In this embodiment, when the note text in the draft script text has been recognized, the note text can be directly deleted, and the draft script text after deleting the note text is used as the script text. In some embodiments, the note text can also be converted into body text, so as to remove the note text without losing the information contained in the draft script text. For example, the text in the draft script text "Common antipyretic analgesics (paracetamol, ibuprofen, and aspirin)" can be converted into "Common antipyretic analgesics, for example, paracetamol, ibuprofen, and aspirin".

[0053] In some cases, to enhance the logic of expression, users may use numbers in the draft script text to indicate the order between each text statement. However, the usage habits of numbers vary among different users, and even the format of the number symbols or number digits used by the same user in the context of a draft script text may be different, which may cause confusion in the voice broadcast of the digital human. Therefore, to improve the dissemination effect of the video, the numbers in the draft script text can be uniformly processed.

[0054] In some embodiments, the video production system can identify the number symbols or number digits in the draft script text; and replace the number symbols or number digits with specified numbered text.

[0055] In this embodiment, the number symbols or number digits can be used to express the order of text statements. Specifically, the types of number symbols or number digits can include Chinese number symbols, English number symbols, Arabic number digits, and Roman number digits, and each type of number symbol or number digit can include multiple formats. Taking Chinese number symbols as an example, it can include "One,...; Two,...; Three,...", "(One),...; (Two),...; (Three),..." and so on.

[0056] In this embodiment, the method of identifying the number symbols or number digits in the draft script text is similar to the method of identifying the note text in the draft script text in the above embodiment, and will not be elaborated here.

[0057] In this embodiment, after identifying all numbered symbols or numbered digits in the line text, all numbered symbols or numbered digits can be replaced with designated numbered characters in a unified format. To improve the voice broadcast effect of the digital human, the designated numbered characters can be Chinese numbered symbols. For example, "one,...; two,...; three,...".

[0058] Step S120: Invoke a large language model based on the long text sentence to obtain multiple reference text sentences corresponding to the long text sentence output by the large language model.

[0059] To improve the video production efficiency, the long text sentences in the line text can be divided using a large language model. At the same time, to overcome the "AI hallucination" problem caused by the large language model, only the division result of the long text sentence output by the large language model can be used as a reference.

[0060] In this embodiment, the reference text sentence can be the division result of the long text sentence output by the large language model. Specifically, the number of reference text sentences can be determined according to the long text sentence division rule input to the large language model when invoking the large language model.

[0061] Step S130: Divide the long text sentence into standard text sentences suitable for the video according to the text of the long text sentence included in the multiple reference text sentences.

[0062] After obtaining the multiple reference text sentences output by the large language model, the long text sentence can be divided with the separation positions of the multiple reference text sentences as the benchmark for dividing the long text sentence.

[0063] In this embodiment, the standard text sentence can be used as the line text of the video.

[0064] In this embodiment, the video production system can respectively match the long text sentence with the multiple reference text sentences according to the word order in the long text sentence to obtain the continuous common text between the reference text sentence and the long text sentence; according to the continuous common text corresponding to the multiple reference text sentences, the long text sentence is divided into multiple text sentences in the word order.

[0065] In this embodiment, the continuous common text can be the maximum number of common characters of the reference text sentence and the long text sentence arranged in the word order of the long text sentence. Specifically, each character in the reference text sentence can be compared with each character in the long text sentence, and when the comparison result is that the characters are the same, the character can be used as a common character. For each reference text sentence, all the common characters are arranged in the word order of the long text sentence to obtain the continuous common text.

[0066] In this embodiment, the continuous common texts respectively corresponding to multiple reference text sentences do not overlap in the character order of the long text sentence, so as to reduce the situation of repeated division for the same characters. Further, according to the sentence division between the reference text sentences where the continuous common texts are located, the division positions between the continuous common texts are correspondingly formed, so as to divide the long text sentence into multiple text sentences based on the division positions.

[0067] In some embodiments, the line text also includes the theme text of the video. Among them, the theme text is the title of the video and needs to be displayed in the video cover frame, while the subtitle text is the line text and needs to be displayed in other video frames except the cover frame.

[0068] The video cover frame is the "first impression" of the video for the audience, and the picture quality of the video cover frame has an important impact on the video dissemination effect. Therefore, in order to improve the picture quality of the video cover frame, compared with the subtitle text of the video, the user has higher requirements for the display effect of the theme text of the video. The video production system can divide the theme text into two standard text sentences, so that the theme text of the video is displayed in two lines in the video cover frame.

[0069] In the embodiment of this specification, by obtaining the long text sentence included in the line text corresponding to the video and invoking the large language model based on the long text sentence, multiple reference text sentences corresponding to the long text sentence output by the large language model are obtained, and then the long text sentence is divided into standard text sentences applicable to the video according to the characters of the long text sentence included in the multiple reference text sentences, so as to generate the line text according to the standard text sentences. Thus, the division position of the long text sentence in the video line text is determined by using the large language model, and then the long text sentence is divided according to the division position, so as to reduce the influence of the "AI hallucination" problem on the video production efficiency on the basis of using the large language model to improve the video production efficiency.

[0070] In this embodiment, the theme text is used to express the theme content of the video. If the number of divided lines is too many, it will affect the user's viewing experience. Therefore, when the number of characters of the theme text is large, the theme text is divided into two standard text sentences. Moreover, as the theme text is usually restricted in terms of the number of characters when formulated, the theme text usually does not include too many characters, so that it can be divided into two standard text sentences.

[0071] In some cases, dividing a long text statement into multiple text statements based on multiple reference text statements may result in some short text statements with too few words in the multiple text statements, which cannot meet the word count requirement of the standard text statement. At this time, it is necessary to connect the short text statement with the adjacent text statement so that the multiple text statements obtained by splicing can meet the word count requirement of the standard text statement and improve the display effect of the line text in the video frame.

[0072] In some embodiments, when there are short text statements with a word count less than the specified word count threshold range in the multiple text statements divided by the video production system, the video production system can calculate the total word count of the text statements below the short text statement in the multiple text statements divided according to the text order of the long text statement and the word count of the short text statement; when the total word count is within the specified word count threshold range, connect the short text statement with the text statement below it into a standard text statement.

[0073] For example, the specified word count threshold range can be not less than 10 words and not more than 20 words. When the word count of a text statement is less than 10 words, it will be recognized as a short text statement. For example, the text statements sequentially include "The weather is really nice today" and "It's suitable to go for a walk". "The weather is really nice today" contains 6 words and belongs to a short text statement. "It's suitable to go for a walk" contains 6 words and also belongs to a short text statement. The total word count of these two short text statements is 12 words, which is within the specified word count threshold range. Thus, these two short text statements can be connected into a standard text statement "The weather is really nice today It's suitable to go for a walk". This standard text statement is suitable for display in the video frame.

[0074] In some cases, the total word count of the short text statement and the text statement below it may exceed the specified word count threshold range, that is, the text statement obtained by connecting the short text statement with the text statement below it may be a long text statement. At this time, the division method of the long text statement in the above-mentioned embodiments can be referred to, and the text statement obtained by connecting the short text statement with the text statement below it can be divided by using a large language model.

[0075] In some embodiments, when there are short text statements with the number of characters less than the specified character threshold range among the multiple text statements divided by the video production system, the video production system calculates the total number of characters of the text statements below the short text statement among the multiple divided text statements in the literal order of the long text statement; when the total number of characters exceeds the specified character threshold range, the large language model is called based on the text statements below the short text statement, and the large language model is instructed to output multiple sub - text statements corresponding to the text statements below the short text statement; in the literal order of the text statements below the short text statement, the first - ranked sub - text statement is connected to the end of the short text statement; wherein, when the number of characters of the connected text statement belongs to the specified character threshold range, the connected text statement is used as the standard text statement.

[0076] For example, the specified character threshold range can be not less than 5 characters and not more than 13 characters. The long text statement is "Taking vitamin D daily can effectively enhance human immunity and reduce the probability of getting sick". The divided text statements include "Taking vitamin D daily", "can effectively", and "enhance human immunity and reduce the probability of getting sick". Among them, the text statement "Taking vitamin D daily" contains 8 characters and belongs to the standard text statement, "can effectively" contains 3 characters and belongs to the short text statement, and "enhance human immunity and reduce the probability of getting sick" contains 13 characters and belongs to the standard text statement. The total number of characters of the short text statement "can effectively" and the text statement "enhance human immunity and reduce the probability of getting sick" is 16 characters. At this time, a prompt instruction can be reconstructed based on "enhance human immunity and reduce the probability of getting sick" to call the large language model, and "enhance human immunity and reduce the probability of getting sick" can be further divided, for example, into two sub - text statements "enhance human immunity" and "reduce the probability of getting sick". At this time, the sub - text statement "enhance human immunity" is connected to the end of the short text statement "can effectively", and the resulting "can effectively enhance human immunity" contains 10 characters and belongs to the standard text statement. The sub - text statement "reduce the probability of getting sick" contains 6 characters and can also be used as a standard text statement.

[0077] In fields with strong professionalism such as medical and health, the update and iteration speed of professional terms is relatively fast, while the training of the large language model has lag. Therefore, in some cases, the large language model is difficult to recognize the latest professional terms in the long text statement and only divides them as ordinary words. At this time, the characters forming a professional term may be scattered in different reference text statements, resulting in a reduction in the reference value of the reference text statements and making it impossible to directly divide the long text statement according to the reference text statements. To reduce the occurrence of the above - mentioned situation, the reference value of the multiple reference text statements output by the large language model can be improved by means of a professional term library.

[0078] In some embodiments, the video production system can match multiple reference text statements with the specified phrases in the specified phrase set; in the case where it is determined that the specified phrases in the specified phrase set are divided into multiple reference text statements, the large language model is called again based on the long text statement, and multiple reference text statements corresponding to the long text statement output by the large language model are obtained.

[0079] In this embodiment, the specified phrase set can be a set of phrases with clear meanings in a specific technical field. For example, for the medical and health field, the specified phrase set can be a set of medical terms; for the field of animals and plants, the specified phrase set can be a set of species scientific names.

[0080] For example, "Taking vitamin D daily can effectively enhance human immunity and reduce the chance of getting sick" is divided into "Taking vitamin daily", "D can effectively enhance human immunity", and "Reduce the chance of getting sick". Among them, the medical term "vitamin D" is divided into "vitamin" and "D" and belongs to different text statements. The video production system can reconstruct the prompt instruction to call the large language model and re-divide to obtain multiple reference text statements. For example, "Taking vitamin D daily", "Can effectively enhance human immunity", and "Reduce the chance of getting sick".

[0081] Please refer to Figure 4 One embodiment of this specification can also provide a video production system, and the video production system can include the following modules.

[0082] An acquisition module; used to acquire the line text corresponding to the video; wherein, the line text includes multiple text statements; the text statements include long text statements and standard text statements; the number of words in the standard text statement is within the specified word count threshold range; the number of words in the long text statement exceeds the specified word count threshold range.

[0083] A call module; used to call the large language model based on the long text statement, and obtain multiple reference text statements corresponding to the long text statement output by the large language model.

[0084] A division module; used to divide the long text statement into standard text statements applicable to the video according to the text of the long text statement included in the multiple reference text statements; wherein, the standard text statement is used as the subtitle of the video.

[0085] Regarding the specific functions and effects achieved by the video production system, reference can be made to other embodiments of this specification for explanation, and details will not be repeated here. Each module in the video production system can be implemented in whole or in part through software, hardware, and their combinations. Each module can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0086] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer executes the processing method of the line text in any of the above embodiments.

[0087] The embodiments of this specification also provide a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the processing method of the line text in any of the above embodiments.

[0088] Please refer to Figure 5 . The embodiments of this specification can provide a computer device, which includes: a memory, and one or more processors communicatively connected to the memory; instructions executable by the one or more processors are stored in the memory. When the instructions are executed by the one or more processors, the one or more processors implement the processing method of the line text in any of the above embodiments.

[0089] In some embodiments, the computer device may include a processor, a non-volatile storage medium, an internal memory, a communication interface, a display device, and an input device connected by a system bus. The non-volatile storage medium may store an operating system and related computer programs.

[0090] The user information or user account information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, etc.) involved in multiple embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws and regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0091] It can be understood that the specific examples in this article are only to help those skilled in the art better understand the embodiments of this specification, rather than limiting the scope of the present invention.

[0092] It can be understood that in various embodiments of this specification, the magnitudes of the numbers of the various processes do not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this specification.

[0093] It can be understood that the various embodiments described in this specification can be implemented alone or in combination, and the embodiments of this specification do not limit this.

[0094] Unless otherwise specified, all technical and scientific terms used in the embodiments of this specification have the same meanings as those commonly understood by those skilled in the technical field of this specification. The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the scope of this specification. The term "and / or" used in the embodiments of this specification includes any and all combinations of one or more of the related listed items. The singular forms "a", "above", and "the" used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0095] It can be understood that the processor in the embodiments of this specification can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this specification can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0096] It can be understood that the memory in the embodiments of this specification can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but not limited to, these and any other suitable types of memories.

[0097] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this specification.

[0098] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0099] In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.

[0100] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0101] In addition, the functional units in each embodiment of this specification can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0102] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions in this specification, in essence, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this specification. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0103] As described above, the above are only specific embodiments of this specification, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed in this specification can easily think of changes or substitutions, which should all be covered within the protection scope of this specification. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for processing medical video dialogue text, characterized in that: include: Acquire a dialogue text corresponding to the medical video; wherein the dialogue text includes a plurality of text sentences; the text sentences include long text sentences and standard text sentences; wherein the number of words in the standard text sentences is within a specified word count threshold range; and the number of words in the long text sentences exceeds the specified word count threshold range; Calling a large language model based on the long text sentence to obtain a plurality of reference text sentences output by the large language model and corresponding to the long text sentence; Matching the multiple reference text sentences with designated phrases in a designated phrase set; wherein the designated phrase set is a collection of phrases with clear meanings in the field of medical health; In the case where it is determined that an inseparable designated phrase in the designated phrase set is divided into at least two of the reference text sentences, re-calling the large language model based on the long text sentence to update the multiple reference text sentences; According to the words of the long text sentences included in the updated multiple reference text sentences, the long text sentences are divided into standard text sentences applicable to the medical video; wherein the standard text sentences serve as subtitle texts of the medical video.

2. The method according to claim 1, characterized in that The step of obtaining the dialogue text corresponding to the medical video includes: Recognize the remark text in the draft text of the lines; wherein the remark text is the text used to express the remark content; The remarks in the draft text of the lines are removed to obtain the text of the lines.

3. The method according to claim 2, characterized in that The method further comprises: Identify numbering symbols or numbering numbers in the draft text of the lines; wherein the numbering symbols or numbering numbers are used to express the order of text sentences; Replace the number symbol or number digit with the specified number text.

4. The method according to claim 1, characterized in that The step of dividing the long text sentence into standard text sentences applicable to the medical video according to the words of the long text sentence included in the updated plurality of reference text sentences comprises: According to the order of characters in the long text sentence, the continuous common text between the reference text sentence and the long text sentence is matched respectively with the multiple reference text sentences; wherein the continuous common text corresponding to the multiple reference text sentences respectively does not overlap according to the order of characters in the long text sentence; According to the continuous common texts corresponding to the multiple reference text sentences, the long text sentence is divided into multiple text sentences in the order of words.

5. The method according to claim 4, characterized in that The method further comprises: If there is a short text sentence with fewer words than the specified word count threshold range among the divided multiple text sentences, the word count of the text sentence located below the short text sentence among the divided multiple text sentences and the total word count of the short text sentence are calculated according to the word order of the long text sentence; When the total number of words is within the specified word count threshold range, the short text sentence is connected with the following text sentence to form a standard text sentence.

6. The method according to claim 4, characterized in that The method further comprises: If there is a short text sentence with fewer words than the specified word count threshold range among the divided multiple text sentences, the word count of the text sentence located below the short text sentence among the divided multiple text sentences and the total word count of the short text sentence are calculated according to the word order of the long text sentence; When the total number of words exceeds the specified word count threshold range, calling the large language model based on the text sentence following the short text sentence, and instructing the large language model to output a plurality of sub-text sentences corresponding to the text sentence following the short text sentence; According to the order of words in the text sentences below the short text sentence, the sub-text sentence ranked first is connected to the end of the short text sentence; wherein, when the number of words in the connected text sentence falls within the specified word count threshold range, the connected text sentence is used as a standard text sentence.

7. The method according to claim 1, characterized in that The text sentence includes the subject text of the medical video; The step of dividing the long text sentences into standard text sentences applicable to the medical video according to the words of the long text sentences included in the plurality of reference text sentences comprises: In the case where the subject text is a long text sentence, the subject text is divided into two standard text sentences.

8. A medical video production system, characterized in that: include: An acquisition module is used to acquire a dialogue text corresponding to the medical video; wherein the dialogue text includes a plurality of text sentences; the text sentences include long text sentences and standard text sentences; wherein the number of words in the standard text sentences is within a specified word count threshold range; and the number of words in the long text sentences exceeds the specified word count threshold range; A calling module, configured to call a large language model based on the long text sentence to obtain a plurality of reference text sentences output by the large language model and corresponding to the long text sentence; A matching module, used to match the plurality of reference text sentences with specified phrases in a specified phrase set; wherein the specified phrase set is a collection of phrases with clear meanings in the field of medical health; An updating module, configured to, when it is determined that an inseparable designated phrase in the designated phrase set is divided into at least two of the reference text sentences, re-call the large language model based on the long text sentence to update the plurality of reference text sentences; A division module is used to divide the long text sentence into standard text sentences applicable to the medical video according to the characters of the long text sentence included in the updated multiple reference text sentences; wherein the standard text sentence serves as the subtitle text of the medical video.

9. A computer device, characterized in that: The computer device includes a memory and a processor, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, the method according to any one of claims 1 to 7 can be implemented.

11. A computer program product, characterized in that The computer program product is used to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text simplification with minimal hallucination

    US20240119220A1