Text data extraction and video production method and system based on AI
By performing sharding and semantic analysis of text data and combining image recognition verification, the problem of inaccurate video synthesis in the prior art is solved, and more accurate video synthesis is achieved.
Patent Information
- Application Number
- CN202411857468.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-17
AI Technical Summary
When the prior art uses text data to synthesize videos, the matching pictures are inaccurate, resulting in inaccurate videos.
By obtaining text data, performing sharding and semantic analysis, semantic representation is identified, and then matching image materials are found in the preset image material library, and the text meaning is verified through image recognition to ensure the accuracy of the match.
Through reverse matching verification, the problem of inaccurate keyword matching is avoided and the accuracy of synthetic videos is ensured.
Smart Images

Figure CN120017776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AI-based text data extraction and video production method and system. Background Art
[0002] When users process documents, they are mostly presented in text form, and it is laborious for users to read text. Therefore, text information can be converted into videos. In this way, users can listen to audio and watch video images to understand the information conveyed by the article without the effort of interpreting the text, which can reduce the difficulty of users obtaining information. Or, because the text is long and it takes time for users to read the text, users do not have the energy to read it one by one, so they need to convert the article into a video, quickly understand the information conveyed by the article through the video, and then choose the article they are interested in to read carefully. In addition, because the video presentation form is more diversified, it is easier to attract the user's attention than boring text reading, and users are more willing to read articles in this way.
[0003] In the prior art, keywords are mainly extracted from text data, and for each keyword, a video image matching the keyword is searched in a pre-set image library, and then the video images are synthesized to obtain the target video. However, the images in the image library are also stored by annotating the corresponding keyword information. Since the keywords are used for matching, the searched images may not be so appropriate to the original text information, resulting in an inaccurate synthesized video. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an AI-based text data extraction and video production method and system, aiming to solve the problem of inaccurate video synthesis using text data in the prior art.
[0005] An object of the present invention is to provide a method for text data extraction and video production based on AI, the method comprising: Acquire text data required for video production, segment the text data according to a preset rule to obtain a plurality of text data segments, and perform semantic analysis on the text data segments to identify semantic representations of the text data segments; Using the semantic representation, searching for a plurality of picture materials matching the semantic representation in a preset picture material library; Performing image recognition on the multiple picture materials to identify textual meanings of the picture materials, and using picture materials whose textual meanings match the text data segment as video pictures of the text data segment; After all the video images of the text data are confirmed, the video images are spliced to obtain the video corresponding to the text data.
[0006] Furthermore, in the above-mentioned AI-based text data extraction and video production method, the step of performing image recognition on the multiple picture materials to identify the text meaning of the picture materials, and using the picture materials whose text meaning matches the text data segment as the video picture of the text data segment includes: Segmenting the image material according to a preset rule to obtain a plurality of sub-image materials, and identifying the sub-image materials respectively to determine the text meaning of each sub-image material; The text meaning of each sub-picture material is matched with the text data segment to determine the picture material that matches the text data segment.
[0007] Furthermore, in the above-mentioned AI-based text data extraction and video production method, the step of segmenting the picture material according to a preset rule to obtain a plurality of sub-picture materials includes: Performing element analysis on the picture material to detect a background portion and a plurality of different characteristic portions of the picture material; The different characteristic parts are segmented out, and the different characteristic parts are combined with the background part to form the multiple sub-picture materials.
[0008] Furthermore, in the above-mentioned AI-based text data extraction and video production method, the step of performing element analysis on the picture material to detect the background part and multiple different feature parts of the picture material includes: Gray-scaling the image material to obtain a grayscale image, and performing edge detection on the grayscale image to obtain a preliminary feature area; Pixel points around the feature region are obtained, and the pixel points are used as seed points to perform region growth on the feature region to obtain a final target feature region, and the target feature region is used as a segmentation region of the feature part.
[0009] Furthermore, in the above-mentioned AI-based text data extraction and video production method, the step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed includes: After all the video pictures of the text data are confirmed, scene information of the video pictures is acquired in sequence, and a video picture group consisting of video pictures with similar scene information is determined in sequence; The style parameters of the video picture group are determined according to a preset rule, and the style parameters are pasted to the video pictures in the video picture group.
[0010] Furthermore, in the above-mentioned AI-based text data extraction and video production method, the step of determining the style parameters of the video picture group according to preset rules and pasting the style parameters into the picture materials in the video picture group includes: Obtaining the picture style parameters of the first video picture in the video picture group, copying the picture style parameters, and then pasting the picture style parameters to the remaining video pictures; or.
[0011] An average value of style parameters of the video pictures in the video picture group is obtained, and the average value of the style parameters is determined as the style parameter of the video picture group.
[0012] Furthermore, in the above-mentioned AI-based text data extraction and video production method, after the step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed, the step further includes: Obtaining a style parameter in one of the video picture groups and a text data segment corresponding to a video picture in the video picture group; Extracting the corresponding text data fragment and letter elements and numeric elements of the style parameters to form an encryption sequence; Performing hash operations on the encryption sequences to obtain sub-encryption chains, connecting the sub-encryption chains end to end to form an encryption chain circle, removing one of the encryption elements at two symmetrical positions of the encryption chain circle and expanding the encrypted elements to obtain a characteristic encryption chain; An encryption key is determined according to the characteristic encryption chain and the sub-encryption chain, and the video is stored according to the encryption key.
[0013] Another object of the present invention is to provide an AI-based text data extraction and video production system, the system comprising: An acquisition module is used to acquire text data required for video production, segment the text data according to a preset rule to obtain a plurality of text data segments, and perform semantic analysis on the text data segments to identify semantic representations of the text data segments; A search module, used to use the semantic representation to search for a plurality of picture materials matching the semantic representation in a preset picture material library; A recognition module, configured to perform image recognition on the plurality of picture materials to identify textual meanings of the picture materials, and use picture materials whose textual meanings match the text data segments as video pictures of the text data segments; A splicing module is used to splice the video images after all the video images of the text data are confirmed to obtain the video corresponding to the text data. Another object of the present invention is to provide a readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of any one of the methods described above.
[0014] Another object of the present invention is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0015] The present invention obtains text data that needs to be used for video production, segments the text data according to preset rules to obtain multiple text data segments, performs semantic analysis on the text data segments to identify the semantic representation of the text data segments; uses the semantic representation to find multiple picture materials that match the semantic representation in a preset picture material library; performs image recognition on multiple picture materials to identify the text meaning of the picture materials, and uses the picture materials whose text meaning matches the text data segment as the video picture of the text data segment; after all video pictures of the text data are confirmed, the video pictures are spliced to obtain the video corresponding to the text data, and after finding the corresponding picture material, the text meaning obtained by image recognition is used to perform reverse matching verification with the text data, thereby avoiding the problem of inaccurate pictures matched by only using keywords, and ensuring the accuracy of the subsequently synthesized video. The problem of inaccurate video synthesis using text data in the prior art is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flow chart of a method for extracting text data and making videos based on AI in the first embodiment of the present invention; Figure 2 This is a structural block diagram of an AI-based text data extraction and video production system in the third embodiment of the present invention.
[0017] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0018] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0019] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0021] Embodiment 1 See also Figure 1 , shown is an AI-based text data extraction and video production method in the first embodiment of the present invention, and the method includes steps S10 to S13.
[0022] Step S10, obtaining text data required for video production, segmenting the text data according to a preset rule to obtain a plurality of text data segments, and performing semantic analysis on the text data segments to identify semantic representations of the text data segments.
[0023] The source of text data may be articles, news, documents, etc. When searching for images, the text data is processed and analyzed to facilitate the subsequent video presentation. Specifically, the text data is segmented to divide the text data into multiple independent text data segments, each of which contains a relatively complete content. The text data segments are then semantically analyzed to identify the semantic representation of the text data, that is, the semantic content. Specifically, the BERT model can be used to understand the deep semantics implied in the text.
[0024] Step S11, using the semantic representation to search for a plurality of picture materials matching the semantic representation in a preset picture material library.
[0025] Specifically, each image in the database has a corresponding label, such as the scene, theme, and other keywords of the image. Similar image materials are found by matching and searching through the semantic meaning of the text data. Generally, multiple related image materials can be found through this rough search method.
[0026] Step S12, performing image recognition on the multiple picture materials to identify the text meanings of the picture materials, and using the picture materials whose text meanings match the text data segment as the video pictures of the text data segment.
[0027] Among them, in order to ensure the accuracy of the image matched by the text data, a reverse matching verification method is used for the second matching. Specifically, image recognition is performed on the found alternative picture materials to obtain the text meaning of the picture materials, and the text meaning of each picture is matched with the text data fragment respectively, so as to find the target picture material with the highest matching degree, and use the picture material as the video picture of the text data fragment.
[0028] Specifically, when confirming the match between text data and image material, similarity is used for matching judgment, wherein the visual transformer and convolutional neural network can be used to identify the text meaning in the image, and the similarity between the text meaning and the vector of the text data segment is used for matching. Exemplarily, cosine similarity or Euclidean distance can be used to measure the similarity between the two, so as to find image material with a high degree of matching.
[0029] In addition, in some optional embodiments of the present invention, in order to further improve the accuracy of matching, the step of performing image recognition on the multiple picture materials to identify the text meaning of the picture materials, and using the picture material whose text meaning matches the text data segment as the video picture of the text data segment includes: Segmenting the image material according to a preset rule to obtain a plurality of sub-image materials, and identifying the sub-image materials respectively to determine the text meaning of each sub-image material; The text meaning of each sub-picture material is matched with the text data segment to determine the picture material that matches the text data segment.
[0030] The basic idea of image segmentation is to divide an image into multiple sub-regions or segments, each of which can be regarded as a local area of the image. Each segment can represent a part of the image information, which may contain a detail of a specific object, background or scene in the image. For each segment, the image and text can be matched independently, and then a comprehensive judgment can be made. For example, the segmented parts of the image can correspond to the parts of the text data segment, and when multiple image segments are related to the text, the matching results of each segment can be weighted and summarized to obtain the corresponding matching result.
[0031] Specifically, the step of dividing the picture material according to a preset rule to obtain a plurality of sub-picture materials includes: Performing element analysis on the picture material to detect a background portion and a plurality of different characteristic portions of the picture material; The different characteristic parts are segmented out, and the different characteristic parts are combined with the background part to form the multiple sub-picture materials.
[0032] Among them, the image is segmented based on the region. For example, different objects or regions in the image (such as background, people, buildings, trees, etc.) can be used as the division criteria. Specifically, the image is analyzed by elements to detect the background part and multiple different feature parts of the image material, and the different feature parts are segmented out. The segmented feature parts are combined with the background parts to form multiple transformed sub-image materials, so that every detail information in the image can be accurately identified during image recognition.
[0033] In addition, in some optional embodiments of the invention, in order to accurately determine the feature division area, the step of performing element analysis on the picture material to detect the background part and multiple different feature parts of the picture material includes: Gray-scaling the image material to obtain a grayscale image, and performing edge detection on the grayscale image to obtain a preliminary feature area; Pixel points around the feature region are obtained, and the pixel points are used as seed points to perform region growth on the feature region to obtain a final target feature region, and the target feature region is used as a segmentation region of the feature part.
[0034] Among them, how to reasonably divide the image and ensure that each segment contains enough semantic information is a challenge. If the division is improper (for example, important objects in the image are divided into multiple segments), it may affect the accuracy of the matching. Therefore, the image material is grayed to obtain a grayscale image, and the edge detection of the grayscale image is performed to obtain the preliminary feature area, and then the region growing is used to determine the final feature segmentation area. Among them, the goal of the edge detection method is to identify areas in the image where the grayscale changes dramatically. Usually these areas are the boundaries of objects or significant features in the image. By performing edge detection on the image, an edge map can be generated. This map shows where the grayscale changes (i.e., edges) have occurred in the image and where there are no changes.
[0035] However, edge detection itself does not always provide complete segmentation results, especially when the edges are discontinuous, noisy, or the contrast between objects is low. Therefore, using region growing to further refine and expand the edges can effectively avoid the shortcomings of edge detection and obtain more accurate and coherent segmentation results. In specific implementation, some pixels near the edges can be selected as seed points for region growing. Usually, edge points themselves may not be suitable as seeds directly because they may be located at the junction of two different regions. Therefore, selecting pixels near the edge as seeds can improve the accuracy of segmentation. Starting from these seed points, the region is expanded according to grayscale similarity through the region growing method. Since the boundary information has been obtained through edge detection, region growing will be more strongly constrained, thus avoiding mis-segmentation.
[0036] Step S13, after all the video images of the text data are confirmed, the video images are spliced to obtain the video corresponding to the text data.
[0037] After all converted video images are determined, the video images can be sequentially spliced in the order of the text data segments to finally obtain the video corresponding to the text data.
[0038] In summary, in the above embodiment of the present invention, a method for extracting text data and producing videos based on AI is provided, by obtaining text data that needs to be produced for video, slicing the text data according to preset rules to obtain multiple text data segments, performing semantic analysis on the text data segments to identify the semantic representation of the text data segments; using the semantic representation to find multiple picture materials that match the semantic representation in a preset picture material library; performing image recognition on multiple picture materials to identify the text meaning of the picture materials, and using the picture materials whose text meaning matches the text data segments as video pictures of the text data segments; after all video pictures of the text data are confirmed, the video pictures are spliced to obtain the video corresponding to the text data, and after finding the corresponding picture materials, the text meaning obtained by image recognition is used to perform reverse matching verification with the text data, thereby avoiding the problem of inaccurate pictures matched only by keywords, and ensuring the accuracy of the subsequently synthesized video. The problem of inaccurate video synthesis using text data in the prior art is solved.
[0039] Embodiment 2 This embodiment also proposes a method for extracting text data and making videos based on AI. The difference between the method for extracting text data and making videos based on AI in this embodiment and the method for extracting text data and making videos based on AI in the first embodiment is that: The step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed includes: After all the video pictures of the text data are confirmed, scene information of the video pictures is acquired in sequence, and a video picture group consisting of video pictures with similar scene information is determined in sequence; The style parameters of the video picture group are determined according to a preset rule, and the style parameters are pasted to the video pictures in the video picture group.
[0040] Among them, the video pictures generated by text may have inconsistencies in style and color tone between different images, which will affect the splicing effect. Therefore, after obtaining all the video pictures, the scene information of the video pictures is obtained, and the video pictures with similar scenes are grouped into video picture groups. Specifically, the styles of video pictures with similar scenes can be determined to be consistent, such as "sunshine beach" and "beach bathing". Therefore, the pictures of these similar scenes are grouped into video picture groups, and the style parameters of the first video picture in the video picture group can be obtained, and these parameters can be copied and pasted to other video pictures in the group, or the average value of the style parameters of the video pictures in the video picture group can be obtained, and the average value of the style parameters is determined as the style parameters of the video picture group. Specifically, the style parameters can be "saturation", "brightness", "color temperature", etc.
[0041] Furthermore, in some optional embodiments of the present invention, after the step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed, the step further includes: Obtaining a style parameter in one of the video picture groups and a text data segment corresponding to a video picture in the video picture group; Extracting the corresponding text data fragment and letter elements and numeric elements of the style parameters to form an encryption sequence; Performing hash operations on the encryption sequences to obtain sub-encryption chains, connecting the sub-encryption chains end to end to form an encryption chain circle, removing one of the encryption elements at two symmetrical positions of the encryption chain circle and expanding the encrypted elements to obtain a characteristic encryption chain; An encryption key is determined according to the characteristic encryption chain and the sub-encryption chain, and the video is stored according to the encryption key.
[0042] Among them, in many cases, the text data content may include some privacy data. Therefore, after being converted into a video, the video is encrypted and stored. Specifically, the elements generated in the video production process are used as encryption elements. In the specific implementation, the video pictures are grouped before, and the style parameters in one of the video picture groups can be randomly selected as the encrypted digital elements, and the text data fragment corresponding to a random picture in the video picture group is selected as the encrypted letter element, such as the pinyin of the text data. Some of the letter elements can be used to form an encryption sequence with the letter and the digital elements, and the encryption sequence is hashed to obtain a sub-encryption chain. In order to further improve the security of the encryption key, the sub-encryption chain itself is connected end to end to obtain an encryption chain circle, and one of the encryption elements at two symmetrical positions of the encryption chain circle is removed, and one of them is randomly removed, and then the encryption chain circle is expanded and connected in the original order to obtain a feature encryption chain, and the feature encryption chain is combined with the sub-encryption chain to obtain the encryption key for video storage. Specifically, the feature encryption chain and the sub-encryption chain can be directly combined and connected to obtain the encryption key, or some encryption elements can be randomly selected from the sub-encryption chain and added to the feature encryption chain to obtain the encryption key.
[0043] In some optional embodiments of the present invention, the radius value of the sequence circle composed of the first and last links of the encryption sequence can also be obtained, the hash value of the radius value can be determined, and the hash value can be combined with the characteristic encryption chain to obtain the corresponding encryption key.
[0044] In summary, in the above embodiment of the present invention, a method for extracting text data and producing videos based on AI is provided, by obtaining text data that needs to be produced for video, slicing the text data according to preset rules to obtain multiple text data segments, performing semantic analysis on the text data segments to identify the semantic representation of the text data segments; using the semantic representation to find multiple picture materials that match the semantic representation in a preset picture material library; performing image recognition on multiple picture materials to identify the text meaning of the picture materials, and using the picture materials whose text meaning matches the text data segments as video pictures of the text data segments; after all video pictures of the text data are confirmed, the video pictures are spliced to obtain the video corresponding to the text data, and after finding the corresponding picture materials, the text meaning obtained by image recognition is used to perform reverse matching verification with the text data, thereby avoiding the problem of inaccurate pictures matched only by keywords, and ensuring the accuracy of the subsequently synthesized video. The problem of inaccurate video synthesis using text data in the prior art is solved.
[0045] Embodiment 3 See also Figure 2 , shown is an AI-based text data extraction and video production system proposed in the third embodiment of the present invention, the system comprising: The acquisition module 100 is used to acquire text data required for video production, segment the text data according to a preset rule to obtain a plurality of text data segments, and perform semantic analysis on the text data segments to identify semantic representations of the text data segments; A search module 200, for searching a preset picture material library for a plurality of picture materials matching the semantic representation by using the semantic representation; The recognition module 300 is used to perform image recognition on the multiple picture materials to identify the text meaning of the picture materials, and use the picture materials whose text meaning matches the text data segment as the video picture of the text data segment; The splicing module 400 is used to splice the video images after all the video images of the text data are confirmed to obtain the video corresponding to the text data.
[0046] Furthermore, in the above-mentioned AI-based text data extraction and video production device, the step of performing image recognition on the multiple picture materials to identify the text meaning of the picture materials, and using the picture materials whose text meaning matches the text data segment as the video picture of the text data segment includes: Segmenting the image material according to a preset rule to obtain a plurality of sub-image materials, and identifying the sub-image materials respectively to determine the text meaning of each sub-image material; The text meaning of each sub-picture material is matched with the text data segment to determine the picture material that matches the text data segment.
[0047] Furthermore, in the above-mentioned AI-based text data extraction and video production device, the step of segmenting the picture material according to a preset rule to obtain a plurality of sub-picture materials includes: Performing element analysis on the picture material to detect a background portion and a plurality of different characteristic portions of the picture material; The different characteristic parts are segmented out, and the different characteristic parts are combined with the background part to form the multiple sub-picture materials.
[0048] Furthermore, in the above-mentioned AI-based text data extraction and video production device, the step of performing element analysis on the picture material to detect the background part and multiple different feature parts of the picture material includes: Gray-scaling the image material to obtain a grayscale image, and performing edge detection on the grayscale image to obtain a preliminary feature area; Pixel points around the feature region are obtained, and the pixel points are used as seed points to perform region growth on the feature region to obtain a final target feature region, and the target feature region is used as a segmentation region of the feature part.
[0049] Furthermore, in the above-mentioned AI-based text data extraction and video production device, the step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed includes: After all the video pictures of the text data are confirmed, scene information of the video pictures is acquired in sequence, and a video picture group consisting of video pictures with similar scene information is determined in sequence; The style parameters of the video picture group are determined according to a preset rule, and the style parameters are pasted to the video pictures in the video picture group.
[0050] Furthermore, in the above-mentioned AI-based text data extraction and video production device, the step of determining the style parameters of the video picture group according to a preset rule and pasting the style parameters to the picture materials in the video picture group includes: Obtaining the picture style parameters of the first video picture in the video picture group, copying the picture style parameters, and then pasting the picture style parameters to the remaining video pictures; or.
[0051] An average value of style parameters of the video pictures in the video picture group is obtained, and the average value of the style parameters is determined as the style parameter of the video picture group.
[0052] Furthermore, in the above-mentioned AI-based text data extraction and video production device, after the step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed, it also includes: Obtaining a style parameter in one of the video picture groups and a text data segment corresponding to a video picture in the video picture group; Extracting the corresponding text data fragment and letter elements and numeric elements of the style parameters to form an encryption sequence; Performing hash operations on the encryption sequences to obtain sub-encryption chains, connecting the sub-encryption chains end to end to form an encryption chain circle, removing one of the encryption elements at two symmetrical positions of the encryption chain circle and expanding the encrypted elements to obtain a characteristic encryption chain; An encryption key is determined according to the characteristic encryption chain and the sub-encryption chain, and the video is stored according to the encryption key.
[0053] The functions or operation steps implemented when the above modules are executed are substantially the same as those in the above method embodiments, and will not be repeated here.
[0054] Embodiment 4 Another aspect of the present invention further provides a readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method described in any one of the above-mentioned embodiments 1 to 2 are implemented.
[0055] Embodiment 5 Another aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in any one of the above-mentioned embodiments 1 to 2 when executing the program.
[0056] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0057] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable storage medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, a "computer-readable storage medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0058] More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable storage medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0059] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0060] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0061] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A text data extraction and video production method based on AI, characterized in that: The method comprises: Acquire text data required for video production, segment the text data according to a preset rule to obtain a plurality of text data segments, and perform semantic analysis on the text data segments to identify semantic representations of the text data segments; Using the semantic representation, searching for a plurality of picture materials matching the semantic representation in a preset picture material library; Performing image recognition on the multiple picture materials to identify textual meanings of the picture materials, and using picture materials whose textual meanings match the text data segment as video pictures of the text data segment; After all the video images of the text data are confirmed, the video images are spliced to obtain the video corresponding to the text data.
2. According to the AI-based text data extraction and video production method of claim 1, it is characterized in that: The step of performing image recognition on the plurality of picture materials to identify the text meaning of the picture materials, and using the picture materials whose text meaning matches the text data segment as the video picture of the text data segment comprises: Segmenting the image material according to a preset rule to obtain a plurality of sub-image materials, and identifying the sub-image materials respectively to determine the text meaning of each sub-image material; The text meaning of each sub-picture material is matched with the text data segment to determine the picture material that matches the text data segment.
3. The AI-based text data extraction and video production method according to claim 2, characterized in that: The step of dividing the picture material according to a preset rule to obtain a plurality of sub-picture materials comprises: Performing element analysis on the picture material to detect a background portion and a plurality of different characteristic portions of the picture material; The different characteristic parts are segmented out, and the different characteristic parts are combined with the background part to form the multiple sub-picture materials.
4. The AI-based text data extraction and video production method according to claim 3 is characterized in that: The step of performing element analysis on the picture material to detect the background part and a plurality of different characteristic parts of the picture material comprises: Gray-scaling the image material to obtain a grayscale image, and performing edge detection on the grayscale image to obtain a preliminary feature area; Pixel points around the feature region are obtained, and the pixel points are used as seed points to perform region growth on the feature region to obtain a final target feature region, and the target feature region is used as a segmentation region of the feature part.
5. The AI-based text data extraction and video production method according to claim 1, characterized in that: The step of splicing the video images to obtain the video corresponding to the text data after all the video images of the text data are confirmed includes: After all the video pictures of the text data are confirmed, scene information of the video pictures is acquired in sequence, and a video picture group consisting of video pictures with similar scene information is determined in sequence; The style parameters of the video picture group are determined according to a preset rule, and the style parameters are pasted to the video pictures in the video picture group.
6. The AI-based text data extraction and video production method according to claim 5, characterized in that: The step of determining the style parameters of the video picture group according to a preset rule and pasting the style parameters to the picture materials in the video picture group comprises: Obtaining the picture style parameters of the first video picture in the video picture group, copying the picture style parameters, and then pasting the picture style parameters to the remaining video pictures; or. An average value of style parameters of the video pictures in the video picture group is obtained, and the average value of the style parameters is determined as the style parameter of the video picture group.
7. The AI-based text data extraction and video production method according to claim 6, characterized in that: After all the video images of the text data are confirmed, the step of splicing the video images to obtain the video corresponding to the text data also includes: Obtaining a style parameter in one of the video picture groups and a text data segment corresponding to a video picture in the video picture group; Extracting the corresponding text data fragment and letter elements and numeric elements of the style parameters to form an encryption sequence; Performing hash operations on the encryption sequences to obtain sub-encryption chains, connecting the sub-encryption chains end to end to form an encryption chain circle, removing one of the encryption elements at two symmetrical positions of the encryption chain circle and expanding the encrypted elements to obtain a characteristic encryption chain; An encryption key is determined according to the characteristic encryption chain and the sub-encryption chain, and the video is stored according to the encryption key.
8. An AI-based text data extraction and video production system, characterized in that: The system comprises: An acquisition module is used to acquire text data required for video production, segment the text data according to a preset rule to obtain a plurality of text data segments, and perform semantic analysis on the text data segments to identify semantic representations of the text data segments; A search module, used to use the semantic representation to search for a plurality of picture materials matching the semantic representation in a preset picture material library; A recognition module, configured to perform image recognition on the plurality of picture materials to identify textual meanings of the picture materials, and use picture materials whose textual meanings match the text data segments as video pictures of the text data segments; A splicing module is used to splice the video images after all the video images of the text data are confirmed to obtain the video corresponding to the text data.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the program.
Citation Information
Patent Citations
Video synthesis method and device, computer equipment and storage medium
CN114390217A
Data query method and device based on artificial intelligence and storage medium
CN115328956A
KR20240167091A