Video generation method and device

By using geolocation data to construct a quadtree structure for scene determination, the method enhances video generation efficiency and quality by automating scene planning and ensuring logical transitions.

CN120321467APending Publication Date: 2025-07-15SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410240.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the generation of videos, the reliance on manual sorting or simple chronological order results in the lack of spatial logical coherence in scene switching, which is inefficient and ineffective.

Method used

By obtaining materials with geographical location information, a quad-tree is built, the video scene and switching order are determined based on geographical location information, and the scene planning is optimized based on spatial logic and shooting time information.

Benefits of technology

It improves the automation and intelligence of video scene planning, improves the efficiency and effect of video generation, and ensures the naturalness and accuracy of scene switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321467A_ABST
    Figure CN120321467A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method, and the method comprises the steps: obtaining a material set of a to-be-generated video, and the material set comprises a plurality of materials with geographic position information; determining scenes of the to-be-generated video and a switching sequence of the scenes based on the geographical location information; and generating a video based on the scenes and the switching sequence of the scenes. According to the technical scheme of the embodiment of the invention, the scene in the generated video and the scene switching sequence can be automatically determined, and the automation and intelligence of scene planning in the video are improved, so that the efficiency and effect of video generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of video technology, and in particular, to a video generation method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the popularization of devices such as smartphones and digital cameras, people have taken a large number of photos with geographical location information and hope to conveniently generate video manuscripts with a certain theme from these photos, such as travel videos.

[0003] Currently, when generating these videos, the planning of the scenes in the video mostly relies on manual sorting by humans, or is based on a simple time sequence, or according to a preset fixed template. However, the method of relying on manual sorting by humans is time-consuming and laborious, with low efficiency; while the method based on a simple time sequence or a fixed template often ignores the spatial connection between the photo shooting locations, resulting in a lack of logical coherence in the spatial switching of the generated video and poor generation effects.

[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a video generation method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above technical problems.

[0006] One aspect of embodiments of the present application provides a video generation method, the method comprising: Obtaining a material set for the video to be generated, the material set including a plurality of materials with geographical location information; Determining the scenes of the video to be generated and the switching order of the scenes based on the geographical location information; Generating a video based on the scenes and the switching order of the scenes.

[0007] Optionally, the determining the scenes of the video to be generated and the switching order of the scenes based on the geographical location information includes: Constructing a quadtree of the material set based on the geographical location information; Determining the scenes of the video to be generated and the switching order of the scenes based on the quadtree.

[0008] Optionally, the geographical location information includes geographical location coordinates; Correspondingly, the constructing a quadtree of the material set based on the geographical location information includes: Map the geographical location coordinates to a two-dimensional plane space, and each geographical location coordinate of the material corresponds to a set of two-dimensional plane coordinates; Use the two-dimensional plane space as the root node of a quadtree, and divide the two-dimensional plane space into four sub-regions; After determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material, add it to the target sub-region, and the target sub-region is any one of the divided sub-regions; When the number of materials in the target sub-region reaches the first threshold and does not reach the predetermined subdivision level, further divide the target sub-region according to the quadtree structure, and add the materials included in the target sub-region to the further divided sub-regions according to the corresponding two-dimensional plane coordinates; Return to execute the steps after adding the material to the target sub-region after determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material, until all materials are added to the quadtree.

[0009] Optionally, the material further includes shooting time information; Correspondingly, the determining the scene of the video to be generated and the switching order of the scenes based on the quadtree includes: Divide the materials that meet the target conditions into the same scene to determine all the scenes included in the video to be generated, where the target conditions include at least one of the materials belonging to the same child node or adjacent child nodes of the quadtree, the time difference between the materials being less than the second threshold, and the content relevance between the materials being greater than the third threshold; Determine the spatial logical order between the scenes based on the spatial relationship corresponding to the quadtree; Determine the switching order of the scenes based on the spatial logical order and the shooting time information.

[0010] Optionally, the generating a video based on the scene and the switching order of the scenes includes: Determine the importance weight and the number of materials of each scene, where the importance weight is used to represent the weight of the importance of the scene in the video; Determine the display duration of each scene based on the importance weight and the number of materials; Generate a video based on the scene, the switching order of the scenes, and the display duration.

[0011] Optionally, the method further includes: When receiving an update operation on the material set, update the quadtree based on the update operation, and the update operation includes adding materials, deleting materials, and modifying the material attributes; Correspondingly, determining the scene of the video to be generated and the switching order of the scenes based on the quadtree includes: Determining the scene of the video to be generated and the switching order of the scenes based on the updated quadtree.

[0012] Optionally, generating a video based on the scene and the switching order of the scenes includes: Obtaining target information of corresponding materials in each scene, where the target information includes at least some of address information, shooting time information, title information, and picture content information; Generating a voice commentary text corresponding to each scene based on the target information; Converting the voice commentary text into a voice commentary audio by using a text-to-speech method; Generating a video based on the scene, the switching order of the scenes, and the voice commentary audio.

[0013] Another aspect of the embodiments of the present application provides a video generation device, and the device includes: An acquisition module, configured to acquire a material set of a video to be generated, where the material set includes multiple materials with geographical location information; A determination module, configured to determine the scene of the video to be generated and the switching order of the scenes based on the geographical location information; A generation module, configured to generate a video based on the scene and the switching order of the scenes.

[0014] Another aspect of the embodiments of the present application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: By obtaining the material set of the video to be generated, determining the scene of the video to be generated and the switching order of the scenes based on the geographical location information of the materials in the material set, and generating the video based on the scene of the video and the switching order of the scenes, the scene and the switching order of the scenes in the video can be automatically determined according to the geographical location information in the materials, effectively improving the automation and intelligence of scene planning in the video, thereby improving the efficiency and effect of video generation. Description of the Drawings

[0018] The drawings exemplarily show embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0019] Figure 1 Schematically shows the flowchart of the video generation method according to Embodiment 1 of the present application; Figure 2 Schematically shows Figure 1 the sub-step flowchart of step S102 in Figure 3 Schematically shows Figure 2 the sub-step flowchart of step S200 in Figure 4 Schematically shows an example diagram of quadtree construction; Figure 5 Schematically shows Figure 2 the sub-step flowchart of step S202 in Figure 6 Schematically shows an example diagram of the scene switching order; Figure 7 Schematically shows Figure 1 the sub-step flowchart of step S104 in Figure 8 Schematically shows Figure 1 another sub-step flowchart of step S104 in Figure 9 Schematically shows the block diagram of the video generation device according to Embodiment 2 of the present application; and Figure 10 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application. Detailed Embodiments

[0020] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0021] It should be noted that in the embodiments of this application, the descriptions involving "first", "second", etc. are only for descriptive purposes and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0022] In the description of this application, it should be understood that the numerical labels before the steps do not identify the sequence of execution of the steps, but are only used to facilitate the description of this application and distinguish each step. Therefore, it cannot be construed as a limitation to this application.

[0023] First, provide the following explanations of the terms involved in this application: Quadtree: A tree-like data structure used to divide a two-dimensional space. It recursively subdivides the two-dimensional space into four quadrants (child nodes), and each node represents a rectangular area in the space. The root node corresponds to the entire spatial range. After one division, four child nodes are obtained, and each child node can be further recursively divided until a specific stop condition is met.

[0024] Video scene: In video content, a relatively independent segment divided according to certain logic, theme, or spatio-temporal relationship. Each scene usually contains a set of related picture contents, jointly expressing specific information or plot.

[0025] Exchangeable image file format (Exif): A file format specifically set for digital camera photos, which can record the attribute information and shooting data of digital photos.

[0026] Text-to-Speech (TTS): Also known as speech synthesis, it is a technology that converts written text into human speech.

[0027] Secondly, to facilitate the understanding of the technical solutions provided in the embodiments of this application by those skilled in the art, the following explains the related technologies: Currently, when users generate videos based on photos, the planning of the scenes in the video mostly relies on manual sorting by humans, or follows a simple chronological order, or is based on preset fixed modules. However, the method of relying on manual sorting by humans is time-consuming and laborious, with low efficiency. Especially when there are a large number of photos, the generation speed is slow. While the methods of following a simple chronological order or based on preset fixed templates often ignore the spatial connection between the photo shooting locations, resulting in a lack of logical coherence in the spatial switching of the generated video and poor video effects.

[0028] Therefore, the embodiments of the present application provide a video generation technical solution. In this technical solution, by determining the scenes in the video and the switching order of the scenes according to the geographical location information in the materials, the scenes in the video and the switching order of the scenes can be automatically determined, improving the automation and intelligence of the scene planning in the video, thereby improving the efficiency and effect of video generation. See the following for details.

[0029] The technical solution of the present application will be introduced through multiple embodiments below. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0030] Embodiment 1 Figure 1 The flowchart of the video generation method according to Embodiment 1 of the present application is schematically shown. It should be noted that the execution subject of the video generation method in the embodiments of the present application can be a client or a server.

[0031] As Figure 1 shown, the video generation method may include steps S100 to S104, where: Step S100: Obtain the material set of the video to be generated, where the material set includes multiple materials with geographical location information.

[0032] Step S102: Determine the scenes in the video to be generated and the switching order of the scenes based on the geographical location information.

[0033] Step S104: Generate a video based on the scenes and the switching order of the scenes.

[0034] The video generation method provided in this embodiment, by obtaining the material set of the video to be generated, determining the scenes in the video to be generated and the switching order of the scenes based on the geographical location information of the materials in the material set, and generating a video based on the scenes and the switching order of the video, can automatically determine the scenes in the video and the switching order of the scenes according to the geographical location information in the materials, effectively improving the automation and intelligence of the scene planning in the video, thereby improving the efficiency of video generation and the efficiency of video generation.

[0035] The following will elaborate in detail on each step in steps S100 to S104 and optional other steps in conjunction with Figure 1 。

[0036] Step S100 ,obtain a material set for the video to be generated, where the material set includes multiple materials with geographical location information.

[0037] Among them, the materials in the material set can include pictures, videos, audios, animations, etc.

[0038] The material set can be obtained after being selected or uploaded by the user of the client. For example, if the user wants to generate a travelogue video using the photos in the local album of the client, then select multiple relevant photos from the album. After the user selects the photos, the client can obtain the material set for the video to be generated based on the selected photos by the user. In the scenario where the server is the execution entity, it can be that after the client uploads the material set, the server obtains the material set. It should be noted that when obtaining the materials in the material set from the client, it is carried out under the authorization of the user.

[0039] The geographical location information of the materials in the material set can be obtained and recorded in the materials through the positioning service of the terminal device during shooting. For example, when the user turns on the positioning service of the mobile phone and takes a photo, the geographical location information of the photo will be recorded in the Exif information of the photo, so that the photo has geographical location information. Optionally, the geographical location information of the materials in the material set can also be added manually later. For example, if some photos do not record geographical location information because the positioning service is not turned on, then the user can manually add geographical location information to these photos. Among them, the geographical location information can include at least one of detailed address, area, city, region, geographical coordinates and other information. Additionally, it should also be noted that the geographical location information of the materials is also recorded under the authorization of the user.

[0040] Step S102 ,determine the scene of the video to be generated and the switching order of the scenes based on the geographical location information.

[0041] Specifically, the materials in the material set can be divided into multiple scenes according to the geographical location information, and then the switching order of the scenes can be determined based on the geographical location information using a preset rule. Among them, the preset rule can be from the center to the periphery, from large to small, from east to west, from south to north, counterclockwise, clockwise, etc. For example, if the materials in the material set include three geographical location information of the city center, suburb and seaside, then the materials corresponding to the city center, suburb and seaside can be divided into separate scenes respectively, and then the switching order of the scenes can be determined as city center → suburb → seaside according to the preset rule from the center to the periphery.

[0042] In an alternative embodiment, such as Figure 2As shown, step S102 may include: Step S200, constructing a quadtree for the material set based on the geographical location information.

[0043] Step S202, determining the scenes of the video to be generated and the switching order of the scenes based on the quadtree.

[0044] Specifically, the geographical location information corresponding to the materials in the material set can be mapped to two-dimensional plane coordinates, the range of the two-dimensional plane space can be determined according to the geographical location information of all materials, and then the two-dimensional plane space is used as the root node of the quadtree, and the two-dimensional plane space is evenly divided into four sub-regions; determine the two-dimensional plane coordinates corresponding to each material in the material set, and add the materials to the sub-regions of the quadtree according to the two-dimensional plane coordinates; if the subdivision condition of the quadtree is reached, continue to subdivide until all materials are added to the quadtree to complete the construction of the quadtree.

[0045] On the basis of completing the construction of the quadtree, some preset rules can be configured according to the quadtree to determine the scenes of the video to be generated and the switching order of the scenes. The preset rules are, for example, dividing the materials belonging to the same sub-region into the same scene, or dividing the materials belonging to adjacent sub-regions into the same scene; or, determining the switching order of each scene according to the sub-regions divided by the quadtree in a clockwise order, such as starting from the upper left sub-region of a region as the first scene, and the upper right, lower right, and lower left sub-regions as the second scene, the third scene, and the fourth scene respectively; or, the switching order of the scenes can be determined according to the hierarchical relationship of the quadtree, such as first generating the scene corresponding to the higher level in the quadtree, and then generating the scene corresponding to the lower level in the quadtree; or, the preset rules can be a combination of the foregoing rules, and the priority of the corresponding rules can also be configured, such as preferentially dividing the materials belonging to the same sub-region into the same scene, and when there are too few materials belonging to the same sub-region, then dividing the materials in the sub-regions adjacent to the sub-region into the same scene, and so on. The specific preset rules can be configured according to actual needs and are not specifically limited here.

[0046] In this embodiment, by constructing a quadtree for the material set based on the geographical location information and determining the scenes of the video to be generated and the switching order of the scenes based on the constructed quadtree, the accuracy of determining the scenes and the switching order of the scenes in the video to be generated can be improved through the construction of the quadtree, thereby improving the quality of the generated video.

[0047] It can be understood that when the geographical location information is relatively rough, it can be used to generate a video with a relatively rough scene sequence. For example, if the geographical location information is a city, and the user has visited several cities within a few days and wants to generate a travelogue video based on the experiences of these days, a relatively rough quadtree can be constructed based on the cities visited by the user within a few days. When the geographical location information is relatively precise, it can be used to generate a video with a relatively precise scene sequence. The following are some exemplary solutions: In an alternative embodiment, the geographical location information includes geographical location coordinates. Correspondingly, in step S200, a quadtree of the material set is constructed based on the geographical location information, as Figure 3 shown, which may include: Step S300, mapping the geographical location coordinates to a two-dimensional plane space.

[0048] Step S302, using the two-dimensional plane space as the root node of the quadtree and dividing the two-dimensional plane space into four sub-regions.

[0049] Step S304, after determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material, adding it to the target sub-region, and the target sub-region is any of the divided sub-regions.

[0050] Step S306, when the number of materials in the target sub-region reaches the first threshold and does not reach the predetermined subdivision level, further subdividing the target sub-region according to the quadtree structure, and adding the materials included in the target sub-region to the further subdivided sub-regions according to the corresponding two-dimensional plane coordinates.

[0051] Step S308, return to execute the steps after determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material and adding it to the target sub-region until all materials are added to the quadtree.

[0052] Specifically, the geographic location coordinates of the materials can be converted into two-dimensional plane coordinates through a projection algorithm (such as Mercator projection), and the range of the two-dimensional plane space can be determined according to the geographic location coordinates of all the materials; then, the obtained two-dimensional plane space is used as the root node of the quadtree, and this root node represents the entire data area; the two-dimensional plane space represented by the root node is evenly divided into four sub-areas; each material in the material set is processed in turn, the two-dimensional plane coordinates after mapping of each material are obtained, and the sub-area where the current material is located is determined as the target sub-area according to the two-dimensional plane coordinates of each material, and the current material is added to the target sub-area; for each divided sub-area (i.e., the target sub-area), check whether the number of materials contained therein reaches the first threshold and whether the sub-area reaches the predetermined subdivision level; in the case where the number of materials contained in the target sub-area reaches the first threshold and does not reach the subdivision level, the target sub-area is further divided according to the structure of the quadtree; after the subdivision is completed, the materials contained in the target sub-area are added to the subdivided sub-areas according to their corresponding two-dimensional plane coordinates; then, return to step S304 and subsequent steps, continue to process the remaining materials in the material set, and continuously repeat these processes until all the materials in the material set are added to the quadtree. Optionally, when the number of materials contained in the target sub-area does not reach the first threshold, it can be directly added without other processing; in the case where the number of materials contained in the target sub-area reaches the first threshold and has reached the predetermined subdivision level, the target sub-area may not be further refined, and when the remaining other materials fall into the target sub-area, they can be added to the target sub-area. Among them, the first threshold and the predetermined subdivision level can be set according to actual needs, and no specific limitation is made here. For example, the first threshold can be 5, and the predetermined subdivision level can be four levels. If a certain target sub-area A contains 5 materials and the current subdivision level is three levels, it is determined that the number of materials reaches the first threshold and does not reach the predetermined subdivision level, then the target sub-area is further divided, and at the same time, the original 5 materials in the target sub-area are added to the subdivided sub-areas.

[0053] Please refer to Figure 4 , which is a schematic diagram of quadtree construction. As shown in the figure, the root node is divided into sub-area 1, sub-area 2, sub-area 3, and sub-area 4. Sub-area 1 is divided into sub-area 1.1, sub-area 1.2, and sub-area 1.3. Photo 1, Photo 3, and Photo 5 are finally added to sub-area 1.2 according to the mapped two-dimensional plane coordinates; Sub-area 2 is divided into sub-area 2.1 and sub-area 2.2. Photo 7 and Photo 8 are finally added to sub-area 2.2 according to the mapped two-dimensional plane coordinates.

[0054] In this embodiment, a quadtree is constructed by mapping the geographical location coordinates of the materials to a two-dimensional plane space. The sub-region where the material is located is determined according to the two-dimensional plane coordinates corresponding to the material and added to the corresponding sub-region. The sub-regions that meet the requirements are further subdivided according to the material quantity threshold of the sub-region and the predetermined subdivision level, and these processes are continuously repeated until all materials are added to the quadtree, so that the quadtree can be effectively constructed according to the geographical location coordinates of the material set, thereby facilitating the subsequent determination of the spatial relationship between each material based on the quadtree.

[0055] In an alternative embodiment, the material further includes shooting time information. Correspondingly, in step S202, based on the quadtree, the scenes of the video to be generated and the switching order of the scenes are determined, as Figure 5 shown, which may include: Step S400, dividing the materials that meet the target conditions into the same scene to determine all the scenes included in the video to be generated, where the target conditions include at least one of the materials belonging to the same child node or adjacent child nodes of the quadtree, the time difference between the materials being less than the second threshold, and the correlation of the picture content between the materials being greater than the third threshold.

[0056] Step S402, determining the spatial logical order between the scenes based on the spatial relationship corresponding to the quadtree.

[0057] Step S404, determining the switching order of the scenes based on the spatial logical order and the shooting time information.

[0058] Specifically, each material in the material set can be judged to determine whether it meets the target conditions, and the materials that meet the target conditions are divided into the same scene. Among them, if the target condition includes the materials belonging to the same child node or adjacent child nodes of the quadtree, the quadtree can be used for judgment, and the materials belonging to the same child node or adjacent child nodes of the quadtree are divided into the same scene; if the target condition includes the time difference between the materials being less than the second threshold, the time difference between each two materials can be calculated, and the materials with a time difference less than the second threshold are divided into the same scene; if the target condition includes the correlation of the picture content between the materials being greater than the third threshold, the correlation of the picture content between each two materials can be calculated. For example, the similarity between each two materials can be calculated through image recognition or feature extraction technology, the correlation is determined according to the similarity, and when the correlation is greater than the third threshold, it is determined that the target conditions are met, and the corresponding materials are divided into the same scene.

[0059] After dividing the scenes, the spatial relationship of each scene can be analyzed according to the structure of the quadtree to determine the spatial logical order between the scenes, such as adjacent relationship, inclusion relationship, from left to right, from top to bottom, clockwise or counterclockwise, etc.; after determining the spatial logical order between the scenes, the switching order of the scenes can be finally determined by combining the shooting time information of the materials. Since the time order and spatial order of the materials do not directly correspond, the priorities of the time order and spatial order can be set. For example, the scenes can be switched in the order of the shooting time information first, and at the same time, the spatial logical order between the scenes should be followed as much as possible. Or, give priority to the spatial logical order. When the spatial logical order cannot determine the sequence of the scenes, determine the switching order of the scenes according to the sequence of the shooting time information. Finally, comprehensively consider the spatial logical order and the time order corresponding to the shooting time information to determine the final switching order of the scenes in the video to be generated.

[0060] Please refer to Figure 6 , which is an example diagram of the scene switching order. As shown in the figure, the materials in the material set can be divided into three scenes: park tour scene, commercial street tour scene, and square tour scene according to the quadtree structure, and then the switching order of the three scenes is determined by combining the shooting time information of the materials: park tour scene → commercial street tour scene → square tour scene.

[0061] In this embodiment, by dividing the materials that meet the target conditions into the same scene, determining the spatial logical order between the scenes based on the spatial relationship corresponding to the quadtree, and combining the spatial logical order and the shooting time information to determine the switching order of the scenes, the spatial logical order and the time order can be effectively combined to determine the switching order of the scenes in the video to be generated, improving the naturalness and accuracy of the scene switching in the video, thereby improving the quality and viewing effect of the generated video.

[0062] Step S104 , generating a video based on the scenes and the switching order of the scenes.

[0063] Specifically, the sequence of the materials in the scene can be determined according to the shooting time information of the materials in the scene, and the materials of each scene can be arranged in order according to the sequence of the materials; between scene switches, appropriate transition effects, such as fades, slides, and zooms, can be added; at the same time, some background music and sound effects can be added, and finally, a video editing software or tool is used to generate a video based on these elements.

[0064] In an alternative embodiment, in step S104, generating a video based on the scenes and the switching order of the scenes, as Figure 7 shown, may include: Step S500, determine the importance weight and the number of materials for each scene, where the importance weight is used to represent the weight of the importance degree of the scene in the video.

[0065] Step S502, determine the display duration of each scene based on the importance weight and the number of materials.

[0066] Step S504, generate a video based on the scenes, the switching order of the scenes, and the display duration.

[0067] Among them, the importance weight of each scene can be obtained according to the analysis result by analyzing the content and related data of the materials in the scene. Specifically, image recognition and scene recognition can be performed on the materials in the scene. If the materials in the scene contain important landmarks, key events or certain special scenes (such as wedding scenes), a higher importance weight can be given. For example, if the materials in the scene are popular scenic spots, a higher importance weight is given; if the materials in the scene are ordinary street scenes, a lower importance weight is given. Alternatively, the importance weight can be determined by combining the interaction data or marked data of the materials in the scene. If the materials in the scene have user interaction data or user marks, a higher importance weight can be given. For example, if the user has interaction data or user marks such as liking for the materials in the scene, a higher importance weight is given.

[0068] The number of materials for each scene can be directly counted for the number of materials in the scene. After obtaining the importance weight and the number of materials for each scene, the display duration of each scene can be determined according to a preset rule. Among them, the preset rule is used to allocate more display duration to those with higher importance weight and more materials. Specifically, it can be through a certain allocation algorithm (such as a linear weighted allocation algorithm) to allocate the specific display duration according to the importance weight and the number of materials of each scene. Optionally, the proportion of the number of materials in each scene to the total number of materials in the material set can be further calculated to normalize the number of materials in each scene, and then the display duration of each scene can be determined according to the importance weight and the proportion of the number of materials. In addition, when allocating the specific display duration according to the importance weight and the number of materials of each scene, it can be to first determine the duration proportion of each scene according to the importance weight and the number of materials of each scene, and then determine the specific display duration of each scene according to the total duration of the video to be generated and the duration proportion of each scene.

[0069] In this embodiment, by determining the importance weight and the number of materials for each scene, determining the display duration of each scene based on the importance weight and the number of materials, and generating a video based on the scenes, the switching order of the scenes, and the display duration, scenes with higher importance and more materials can be preferentially given longer display duration, so as to ensure the rationality of the generated video rhythm and improve the effect and quality of the generated video.

[0070] In an alternative embodiment, the video generation method in the embodiments of the present application may further include: when an update operation on the material set is received, updating the quadtree based on the update operation, where the update operation includes adding materials, deleting materials, and modifying material attributes; correspondingly, in step S202, determining the scenes of the video to be generated and the switching order of the scenes based on the quadtree may include: determining the scenes of the video to be generated and the switching order of the scenes based on the updated quadtree.

[0071] Specifically, when there are new materials in the material set, the geographical location information of the new materials can be obtained, and two-dimensional plane space mapping is performed according to the geographical location information of the new materials to obtain the two-dimensional plane coordinates corresponding to the new materials, and the new materials are added to the quadtree according to the two-dimensional plane coordinates of the new materials; when adding to a certain target sub-region, if the number of materials included in the target sub-region reaches the first threshold and does not reach the predetermined subdivision level, the target sub-region is further subdivided according to the quadtree structure, and then the materials in the target sub-region are reallocated to the subdivided sub-regions. If there are materials deleted in the material set, the corresponding materials in the quadtree can be deleted. If there are materials with modified attributes in the material set, it can first be determined whether the modified attributes affect the quadtree. If the modified attributes (such as geographical location information attributes) affect the quadtree, the materials are re-added to the quadtree according to the modified attributes.

[0072] After the quadtree is updated, determine the scenes of the video to be generated and the scene switching order again according to the updated quadtree according to the foregoing method; it is also possible to re-determine the display duration of each scene of the video to be generated according to the updated quadtree, etc., so that the generated video adapts to the updated operation.

[0073] In this embodiment, by updating the quadtree based on the update operation when an update operation on the material set is received, and determining the scenes of the video to be generated and the switching order of the scenes based on the updated quadtree, the generation of the video can be updated in a timely manner according to the user's update operation, so that the generated video adapts to the latest material status data.

[0074] In an alternative embodiment, in step S104, generating a video based on the scene and the switching order of the scenes, as Figure 8 shown, may include: Step S600, obtaining the target information of the corresponding materials in each scene, where the target information includes at least some of the address information, shooting time information, title information, and picture content information.

[0075] Step S602, generating a voice commentary text corresponding to each scene based on the target information.

[0076] Step S604: Use the text-to-speech method to convert the voice commentary text into voice commentary audio.

[0077] Step S606: Generate a video based on the scene, the switching order of the scenes, and the voice commentary audio.

[0078] Among them, the address information can be obtained from the geographical location information of the material. For example, the geographical location information of the material can be converted into address information through a geocoding service; the picture content information can be obtained by analyzing the content (such as the main object) of the material through image recognition technology (such as an image model of deep learning).

[0079] After obtaining the target information of the material in each scene, the corresponding voice commentary text can be generated according to these target information using a preset template. Among them, the preset template can be set differently according to different target information. For example, the preset template is "Next, we visited [XX (address information)] (only when the address appears for the first time), [picture content information]"; after a certain material adopts this preset template, the corresponding voice commentary text is: Next, we visited the West Lake in Hangzhou. We stood under the willow trees by the lake... When generating the corresponding voice commentary text according to the preset template, some transitional sentences can be appropriately added to make the commentary text of the entire video coherent and natural. Optionally, the corresponding voice commentary text can also be generated using generative artificial intelligence according to these target information. For example, these target information can be input into a generative artificial intelligence model, and the generative artificial intelligence is used to generate the corresponding voice commentary text.

[0080] After obtaining the voice commentary text corresponding to each scene, the voice commentary text can be converted into voice commentary audio using TTS. Finally, the final video is generated by combining the scene, the switching order of the scenes, and the voice commentary.

[0081] In this embodiment, by obtaining the target information of the corresponding material in each scene, generating the voice commentary text corresponding to each scene based on the target information of the material, using the text-to-speech method to convert the voice commentary text into voice commentary audio, and generating a video based on the scene, the switching order of the scenes, and the voice commentary audio, the generated video can be accompanied by automatically generated voice commentary, improving the effect of the generated video.

[0082] Embodiment 2 Figure 9A block diagram of a video generation device according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 9 shown, the device 700 may include: an acquisition module 710, a construction module 720, a determination module 730, and a generation module 740, where: The acquisition module 710 is configured to acquire a material set of the video to be generated, and the material set includes a plurality of materials with geographical location information; The determination module 720 is configured to determine the scene of the video to be generated and the switching order of the scenes based on the geographical location information; The generation module 730 is configured to generate a video based on the scene and the switching order of the scenes.

[0083] In an alternative embodiment, the determination module 720 is further configured to: Construct a quadtree of the material set based on the geographical location information; Determine the scene of the video to be generated and the switching order of the scenes based on the quadtree.

[0084] In an alternative embodiment, the geographical location information includes geographical location coordinates; Correspondingly, the determination module 720 is further configured to: Map the geographical location coordinates to a two-dimensional plane space, and each geographical location coordinate of the material corresponds to a set of two-dimensional plane coordinates; Use the two-dimensional plane space as the root node of the quadtree, and divide the two-dimensional plane space into four sub-regions; After determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material, add it to the target sub-region, and the target sub-region is any one of the divided sub-regions; In the case where the number of materials in the target sub-region reaches a first threshold and does not reach a predetermined subdivision level, further subdivide the target sub-region according to the quadtree structure, and add the materials included in the target sub-region to the further subdivided sub-regions according to the corresponding two-dimensional plane coordinates; Return to execute the steps of determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material and adding it to the target sub-region and subsequent steps until all materials are added to the quadtree.

[0085] In an alternative embodiment, the material further includes shooting time information; Correspondingly, the determining module 720 is further configured to: Divide the materials that meet the target conditions into the same scene to determine all the scenes included in the video to be generated, where the target conditions include at least one of the materials belonging to the same child node or adjacent child nodes of the quadtree, the time difference between the materials being less than a second threshold, and the relevance of the picture content between the materials being greater than a third threshold; Determine the spatial logical order between the scenes based on the spatial relationship corresponding to the quadtree; Determine the switching order of the scenes based on the spatial logical order and the shooting time information.

[0086] In an alternative embodiment, the generating module 730 is further configured to: Determine the importance weight and the number of materials of each scene, where the importance weight is used to represent the weight of the importance of the scene in the video; Determine the display duration of each scene based on the importance weight and the number of materials; Generate a video based on the scenes, the switching order of the scenes, and the display duration.

[0087] In an alternative embodiment, the apparatus 700 is further configured to: When receiving an update operation on the material set, update the quadtree based on the update operation, where the update operation includes adding materials, deleting materials, and modifying the material attributes; Correspondingly, the determining module 720 is further configured to: Determine the scenes of the video to be generated and the switching order of the scenes based on the updated quadtree.

[0088] In an alternative embodiment, the generating module 730 is further configured to: Obtain the target information of the corresponding materials in each scene, where the target information includes at least some of the address information, shooting time information, title information, and picture content information; Generate a voice commentary text corresponding to each scene based on the target information; Convert the voice commentary text into a voice commentary audio by using a text-to-speech method; Generate a video based on the scenes, the switching order of the scenes, and the voice commentary audio.

[0089] Embodiment III Figure 10Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing the video generation method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 10 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the video generation method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.

[0090] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0091] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0092] It should be noted that Figure 10 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.

[0093] In this embodiment, the video generation method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0094] Embodiment 4 The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the video generation method in the embodiments are implemented.

[0095] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the video generation method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.

[0096] Embodiment Five The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.

[0097] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0098] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.

Claims

1. A video generation method, characterized in that, The method includes: Obtaining a material set of the video to be generated, where the material set includes multiple materials with geographical location information; Determining the scene of the video to be generated and the switching order of the scenes based on the geographical location information; Generating a video based on the scene and the switching order of the scenes.

2. The method according to claim 1, wherein The determining the scene of the video to be generated and the switching order of the scenes based on the geographical location information includes: Constructing a quadtree of the material set based on the geographical location information; Determining the scene of the video to be generated and the switching order of the scenes based on the quadtree.

3. The method according to claim 2, characterized in that, The geographical location information includes geographical location coordinates; Correspondingly, the constructing the quadtree of the material set based on the geographical location information includes: Mapping the geographical location coordinates to a two-dimensional plane space, and each geographical location coordinate of the material corresponds to a set of two-dimensional plane coordinates; Taking the two-dimensional plane space as the root node of the quadtree, and dividing the two-dimensional plane space into four sub-regions; After determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material, adding it to the target sub-region, and the target sub-region is any one of the divided sub-regions; When the number of materials in the target sub-region reaches the first threshold and does not reach the predetermined subdivision level, further dividing the target sub-region according to the quadtree structure, and adding the materials included in the target sub-region to the further divided sub-regions according to the corresponding two-dimensional plane coordinates; Return to execute the steps after adding the material to the target sub-region after determining the target sub-region where the material is located according to the two-dimensional plane coordinates corresponding to the material, until all materials are added to the quadtree.

4. The method according to claim 3, wherein The material further includes shooting time information; Correspondingly, the determining the scene of the video to be generated and the switching order of the scenes based on the quadtree includes: Dividing the materials that meet the target conditions into the same scene to determine all the scenes included in the video to be generated, where the target conditions include at least one of the material belonging to the same child node or adjacent child nodes of the quadtree, the time difference between the materials being less than the second threshold, and the content correlation between the materials being greater than the third threshold; Determining the spatial logical order between the scenes based on the spatial relationship corresponding to the quadtree; Determining the switching order of the scenes based on the spatial logical order and the shooting time information.

5. The method according to claim 1, wherein The generating a video based on the scene and the switching order of the scenes includes: Determining the importance weight and the number of materials of each scene, where the importance weight is used to represent the weight of the importance degree of the scene in the video; Determining the display duration of each scene based on the importance weight and the number of materials; Generating a video based on the scene, the switching order of the scenes, and the display duration.

6. The method according to any one of claims 2-5, characterized in that, The method further includes: When receiving an update operation on the material set, updating the quadtree based on the update operation, and the update operation includes adding materials, deleting materials, and modifying the material attributes; Correspondingly, determining the scene of the video to be generated and the switching order of the scenes based on the quadtree includes: Determining the scene of the video to be generated and the switching order of the scenes based on the updated quadtree.

7. The method according to any one of claims 1 to 5, characterized in that, Generating a video based on the scene and the switching order of the scenes includes: Obtaining the target information of the corresponding materials in each scene, where the target information includes at least some of the address information, shooting time information, title information, and picture content information; Generating the voice commentary text corresponding to each scene based on the target information; Converting the voice commentary text into a voice commentary audio by using a text-to-speech method; Generating a video based on the scene, the switching order of the scenes, and the voice commentary audio.

8. A video generation device, characterized in that, The device includes: An obtaining module, configured to obtain a material set of the video to be generated, where the material set includes a plurality of materials with geographical location information; A determining module, configured to determine the scene of the video to be generated and the switching order of the scenes based on the geographical location information; A generating module, configured to generate a video based on the scene and the switching order of the scenes.

9. A computer device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.