Photo volume generation method and device based on large model

By employing a large model-based approach, multimodal large models and diffusion model networks are used for image content description and feature extraction, automatically generating photo rolls. This solves the problem of low efficiency in traditional methods, achieving efficient and accurate image classification and deduplication, and generating high-quality photo rolls.

CN121636739APending Publication Date: 2026-03-10BEIJING HISIGN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Traditional methods for generating photo rolls require a lot of manual work, are inefficient, and cannot meet the needs of rapid analysis and processing of image data in the era of big data. Furthermore, the resolution of the cropped image may not meet the requirements.

Method used

A large model-based approach is adopted, which uses multimodal large models and large language models to describe and classify image content, and combines diffusion model networks for feature extraction and deduplication to automatically generate photo rolls.

Benefits of technology

It achieves end-to-end image classification and automatic photo roll generation, ensuring accurate image classification and deduplication. The generated photo rolls have high quality consistency, strong resolution adaptability, reduce human error, and improve work efficiency and the practicality of photo rolls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636739A_ABST
    Figure CN121636739A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a photo volume generation method and device based on a large model, and the method comprises the steps: describing the contents of all reconnaissance images in a preset range based on a multi-modal large model according to a first cue word, and generating a detail title of each reconnaissance image; carrying out character analysis on the detail title of each survey image according to a second cue word based on a large language model to obtain a spatial hierarchy corresponding to the detail title; according to the investigation images of different spatial levels, feature extraction is carried out on the investigation images based on a diffusion model network, and duplicate removal is carried out on the same view angle images of the same spatial level; on the basis of a diffusion model network, sequentially segmenting each survey image from the survey image set of the uppermost spatial hierarchy, searching for the most similar survey image in the next spatial hierarchy, and performing spatial matching; based on a result of the spatial matching, a photo volume from a complete spatial hierarchy is generated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present document relates to the field of image recognition, and in particular to a photo album generation method and device based on a large model. BACKGROUND

[0002] Photo album generation is a method of analyzing the spatial structure of on-site investigation images and arranging existing images in order according to spatial dependence. The purpose is to classify images according to the size of the space they contain into panoramic, orientation, overview, and detailed images containing different spatial scales, and the latter must be contained in the spatial range shown by the former, so as to obtain a series of ordered images of on-site investigation results from macro to detail. The traditional photo album generation scheme belongs to a subfield of image recognition, involving the recognition of photo spatial range and the analysis of inclusion relationship.

[0003] In the traditional on-site investigation workflow, a large number of global, local, and detail images are often taken on site, and in the subsequent collation stage, images corresponding to different spatial scales are selected manually and combined into a series of spatial scale ordered image sequences to form a composite required photo album. In this traditional process, the production of a photo album requires manual quality screening, type classification, and hierarchical division of a large number of on-site pictures, and the overall process of generating a photo album is relatively inefficient and requires a large amount of manpower, making it difficult to meet the rapid analysis and processing of large amounts of image data in the big data era.

[0004] To solve the above problems, in the prior art, the following technical scheme is proposed: analyzing a spatial structure diagram of a specified place to extract compartment information of at least one compartment in the specified place; determining an investigation area of the at least one compartment according to the compartment information of the at least one compartment; selecting at least one point and a corresponding angle of the at least one point in the investigation area; and performing a screenshot operation on a panoramic image of the specified place based on the at least one point and the corresponding angle to obtain an overview image of the specified place. Thus, an overview image can be generated based on spatial structure information of the specified place.

[0005] In the above technical scheme, on the one hand, human resources are needed to analyze and select the area of the panoramic image in the first step, and on the other hand, only the overview image can be taken from the panoramic image, which has the disadvantage that the resolution of the taken image may not meet the requirements, and the potential overview images taken on site are not fully utilized. SUMMARY

[0006] The purpose of the present application is to provide a photo album generation method and device based on a large model to solve the above problems in the prior art.

[0007] The present application provides a photo album generation method based on a large model, comprising: The multi-modal large model is used to describe the content of all survey images within a predetermined range according to a first prompt word, and a detail title of each survey image is generated. The large language model is used to perform text analysis on the detail title of each survey image according to a second prompt word, and a spatial level corresponding to the detail title is obtained. According to the survey images of different spatial levels, the diffusion model network is used to extract features of the survey images, and the same perspective images of the same spatial level are de-duplicated. Based on the diffusion model network, starting from the survey image set of the uppermost spatial level, each survey image is sequentially segmented, the most similar survey image in the next spatial level is found, and spatial matching is performed. Based on the result of spatial matching, a photo roll from a complete spatial level is generated.

[0008] The application provides a photo roll generation device based on a large model, comprising: The title description module is used to describe the content of all survey images within a predetermined range according to a first prompt word based on a multi-modal large model, and a detail title of each survey image is generated. The spatial level module is used to perform text analysis on the detail title of each survey image according to a second prompt word based on a large language model, and a spatial level corresponding to the detail title is obtained. The feature extraction module is used to extract features of the survey images based on a diffusion model network according to the survey images of different spatial levels, and the same perspective images of the same spatial level are de-duplicated. The spatial matching module is used to sequentially segment each survey image starting from the survey image set of the uppermost spatial level based on the diffusion model network, find the most similar survey image in the next spatial level, and perform spatial matching. The generation module is used to generate a photo roll from a complete spatial level based on the result of spatial matching.

[0009] The application also provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is executed by the processor to implement the steps of the photo roll generation method based on a large model.

[0010] The application also provides a computer readable storage medium having an information transmission implementation program stored thereon, wherein the program is executed by a processor to implement the steps of the photo roll generation method based on a large model.

[0011] By introducing the full-automatic process based on the large model, the end-to-end image classification, deduplication and automatic generation of photo albums are realized, the accurate classification of images is ensured, the appearance of repeated images is effectively avoided, and the quality and consistency of the photo album are ensured. The containing relationship between images can be automatically analyzed, and the structured photo album with spatial dependency is automatically generated. The image resolution can be automatically adjusted according to actual needs, and the image resolution in the generated photo album meets the requirements. Not only the possibility of manual error is reduced, but also the generation process of the photo album is more stable and reliable. The present application can efficiently process a large amount of image data and quickly generate high-quality photo albums.7. The burden of the user is greatly reduced, and the work efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the one or more embodiments of the present specification or the prior art, the drawings needed to be used in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present specification, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] Figure 1 is a flowchart of the photo album generation method based on the large model of the embodiment of the present application; Figure 2 is a schematic diagram of image title generation of the embodiment of the present application; Figure 3 is a schematic diagram of title classification of the embodiment of the present application; Figure 4 is a schematic diagram of the photo album generation device of the large model of the embodiment of the present application; Figure 5 is a schematic diagram of an electronic device of the embodiment of the present application. DETAILED DESCRIPTION

[0014] In order to make the person skilled in the art better understand the technical solutions in the one or more embodiments of the present specification, the technical solutions in the one or more embodiments of the present specification will be described clearly and completely below in conjunction with the drawings in the one or more embodiments of the present specification. Obviously, the described embodiments are only some embodiments of the present specification, not all embodiments. Based on the one or more embodiments of the present specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present document.

[0015] Method embodiment According to the embodiment of the present application, a photo album generation method based on a large model is provided, Figure 1is a flowchart of a large model-based photo album generation method of an embodiment of the present application, as shown in Figure 1 The large model-based photo album generation method according to an embodiment of the present application specifically includes the following steps: Step S101, based on a multi-modal large model, describing the content of all survey images within a predetermined range according to a first prompt word, generating a detailed title for each survey image; specifically including: Obtain the prompt word, wherein the prompt word includes a description of the spatial type of the content shown in the image, a description of the environmental elements of the object shown in the image, a description of the possible object scale, and a description of the spatial depth of the image; Based on the multi-modal large model, the content of all survey images within a predetermined range is described using the prompt word, and the set of all survey images of a certain survey is , then the image content description of any survey image is: Formula 1; Wherein, is the image content description of the image Under the conditions of given multi-modal large model prompt word , generation temperature control parameter and generation sampling parameter , the image content description obtained by sampling.

[0016] Step S102, based on a large language model, text analysis of the detailed title of each survey image according to a second prompt word, obtaining a spatial level corresponding to the detailed title; specifically including: Based on the large language model, according to the given second prompt word, generation temperature control parameter, and sampling parameter, the spatial level to which each survey image belongs is judged, wherein the second prompt word includes all classification types and the definition of each type, and the understanding of the large language model for the classification type to be classified is enhanced through examples and the required output format is specified.

[0017] The spatial level specifically includes: panorama, orientation, overview, and detailed item.

[0018] Step S103, according to survey images of different spatial levels, based on a diffusion model network, extracting features of the survey images, and removing the same perspective images of the same spatial level; specifically including: For all survey images in each spatial level, first feature extraction is performed by a pre-trained image feature extraction network; on this basis, the similarity between any two features is calculated to construct a similarity matrix between all pairs of survey images in the spatial level; based on a pre-set similarity threshold, the survey images are clustered and grouped to obtain the grouping of all survey images in the image subset classified based on spatial scale according to similarity; after completing the clustering and grouping, the survey image with the highest similarity is selected as the representative survey image under the same perspective in each group through selection and group image evaluation.

[0019] Step S104, based on the diffusion model network, starting from the survey image set of the top spatial level, each survey image is sequentially segmented to find the most similar survey image in the next spatial level and perform spatial matching; specifically including: For any two image sets of upper and lower levels, first the survey images of the upper level with a larger spatial range are regionally divided into several sub-regions with certain overlapping intervals, and the divided sub-regions include several different size levels, and after the sub-regions corresponding to each sub-region are preprocessed and cropped, the pre-trained image feature extraction network is used for feature extraction, and the features and corresponding sub-region attribute information are stored in a dictionary manner; The survey images of the next level with a smaller spatial range are feature-extracted using the image feature extraction network, and for the features corresponding to each survey image in the next level spatial range, the sub-regions in the dictionary of sub-region features of the previous level with a similarity exceeding a threshold are found, and the attribute information of the best matching sub-region is recorded; after finding the candidate set of related sub-region images in the previous level corresponding to each survey image in the next level, the geometric consistency is verified by a feature matching algorithm; the results are further de-duplicated to find the spatial relationship between each image in the next level image set and the corresponding large image in the previous level image set, and finally organized into an associated set.

[0020] Step S105, based on the results of spatial matching, generating a photo roll from the complete spatial level. Specifically including: By traversing each adjacent set from panoramic to specific purpose images, a complete photo relationship set from panoramic to specific purpose images is constructed, and the positioning of the spatial relationship of different level images in the present survey image starting from the panoramic image is finally formed into a photo roll.

[0021] As can be seen from the above description, the embodiment of the present application proposes a method of making full use of all image materials shot in the scene investigation stage, combining a multimodal large model and a language large model to make content description and spatial size judgment on different images, and using an image feature extraction network to deduplicate images in the same level, further analyzing the image inclusion relationship and automatically generating a photo album with spatial dependency and structure. The embodiment of the present application uses the powerful image understanding capability of the multimodal large model to analyze and understand the content of all images obtained in the scene investigation, thereby generating detailed title content corresponding to each image; and uses the natural language analysis capability of the large language model to analyze and understand the title content describing a certain image in detail, thereby correctly classifying the spatial level to which the image belongs; further, a pre-trained image feature extraction model is used to extract features of all images belonging to the same spatial level and realize image deduplication; finally, the same image feature extraction model is used to correctly analyze the spatial relationship between images in adjacent levels, and finally realize automatic generation of a photo album of scene investigation results. In the present application, by decoupling the nodes of the image content understanding, classification, deduplication and relationship analysis processes, and automatically analyzing each process, the full-process automation of the photo album generation is realized.

[0022] The above technical solutions of the embodiments of the present application will be described in detail below in combination with the drawings.

[0023] The photo album generation method proposed in the present application includes the following steps: (1) Image title generation. In this step, a multimodal large model is used to generate a detailed title based on a prompt word engineering for each shot scene investigation picture, and the generated title will contain a detailed description of the content contained in the picture. Specifically, let the whole scene investigation image set for a certain investigation be Then, for any image belonging to the set , there is:

[0024] Wherein, is the image In a given multimodal large model prompt word , a generation temperature control parameter and a generation sampling parameter image content description. Since the first step focuses on analyzing the spatial content of the image, the prompt part will focus on enhancing the description and limitation of the image content description related to the spatial distribution of the image. Specifically, the prompt needs to include a description of the spatial type of the content shown in the image, including the types and instances of macro, meso and micro scenes; secondly, it needs to include a description of the environmental elements of the objects shown in the image, including possible buildings, vegetation coverage and spatial boundary features; further, it needs to include a description of the scale of the objects that may exist, such as the possible scale of the typical objects appearing in the image; finally, it needs to include an explanation of the spatial depth of the image. An example of this phase is shown in Figure 2 .

[0025] (2) Image hierarchical classification.

[0026] As Figure 3 shown, through the joint process of steps 1 and 2 above, the detailed understanding of the image content and the classification step can be effectively decoupled, ensuring more accurate classification of the image. At this time, the classification of the entire set of images of a survey is completed, i.e., exists and makes. After completing the classification of the images, the images in each class will be further de-duplicated.

[0027] (3) Image de-duplication. After the above classification is completed, each class contains different images of corresponding spatial scales. In any set of images of a class, some images may be taken from different angles, while others are almost identical pictures taken from the same perspective. In order to make a complete photo album, the images in the same classification level need to be de-duplicated, i.e., only one image is retained under any perspective. In the embodiment of the present application, a pre-trained image feature extraction network AlexNet is used to extract features of all images in the same level, and the similarity between the features of any two images is calculated. Specifically, for all images in each classification level, first, feature extraction is performed through the pre-trained image feature extraction network; on this basis, the similarity between any two extracted features is calculated, thereby constructing a similarity matrix between all image pairs in the classification level; further, based on a pre-set similarity threshold, the images are clustered and grouped, i.e., the grouping of all images in the image subset classified based on spatial scales according to similarity is obtained, thereby ensuring that the images in each group are repeated photo groups taken from the same perspective with high similarity; after clustering and grouping is completed, the image with the highest similarity is selected as the representative image under the same perspective in each group. Through the above process, the present application further de-duplicates the images in the same spatial scale level, ensuring that a complete and unique set of images taken from all perspectives under the corresponding spatial scale level is retained. The pseudo code of the algorithm flow in this phase is as follows: Input: Set of images , Similarity threshold , Optional parameters

[0028] Output: Set of deduplicated images

[0029] Phase 1: Initialization of parameters 1.1 Initialize LPIPS model

[0030] 1.2 Initialize empty set of feature vectors

[0031] 1.3 Initialize empty similarity matrix

[0032] Phase 2: Feature extraction 2. For each image : 2.1 Preprocess (scale, normalize, etc.) 2.2 Extract feature vector of using model

[0033] 2.3 Add to the set of feature vectors Phase 3: Similarity computation 3.1 Initialize similarity matrix as a zero matrix of size x 3.2 For each pair of feature vectors in : 3.2.1 Compute LPIPS distance between and

[0034] 3.2.2 Store in the th row and th column of the similarity matrix Phase 4: Grouping of similar images 4.1 Initialize empty set of groups

[0035] 4.2 For each image index : 4.2.1 If has already been assigned to a group, skip​​​ 4.2.2 Create new group

[0036] 4.2.3 For each : If , add to 4.2.4 Add to the set of groups Phase Five: Select representative image 5.1 Initialize empty set of output indices

[0037] 5.2 For each group : 5.2.1 If , add the group unique element to 5.2.2 Else: a. Calculate the centrality score C for each image in b. Select the image index with the highest centrality

[0038] c. Add to Phase Six: Generate result 6. For each index : 6.1 Add to the output set Phase Seven: Return result 7. Return the set of images with duplicates removed

[0039] 7. Return the set of images with duplicates removed ​​​​​​(4) Image association. After the image hierarchical classification and image deduplication in the same level, the application will further match the spatial relationship between images in different spatial levels. Overall, the goal of this stage is to find the corresponding area of a certain image in the upper level for any image in the lower level. In this application, the image feature extraction network used in the previous stage will be reused. First, for any two image sets in the upper and lower levels (such as panorama and orientation, orientation and overview, and overview and detail), the image in the upper level with a larger spatial range is first divided into several sub-regions with a certain overlap interval, and the sub-regions include several different size levels (such as 3*3 size, 5*5 size, 7*7 size, and a series of different size sub-intervals), and after the sub-region corresponding to the sub-region is preprocessed, the pre-trained image feature extraction network is used for feature extraction, and the feature and corresponding sub-region attribute information (including the image to which the sub-region belongs, the interception range of the sub-region, and the like) are stored in the form of a dictionary; second, the image in the lower level with a smaller spatial range is used to extract features using the same feature extraction network; further, for each image in the lower level, the sub-region feature of the upper level image formed in the first step is searched in the dictionary, and the sub-region with a similarity exceeding a threshold value is recorded, and the attribute information of the best matching sub-region is recorded; after finding the corresponding sub-region image candidate set in the upper level for each image in the lower level, the geometric consistency is verified by a feature matching algorithm; finally, by further deduplication of the results, the spatial relationship between each image in the lower level image set and the corresponding large image in the upper level image set is found, and finally organized into an association set. The pseudo code of the algorithm flow in this stage is as follows: Input: upper level image set , lower level image set , similarity threshold

[0040] Output: membership relationship set (lower level and upper level image region correspondence) Stage one: initialization parameters 1.1 Load feature extraction model

[0041] 1.2 Set region division parameters (window size, overlap rate, etc.) 1.3 Feature storage structure BF (upper level image region feature) and SF (lower level image feature) Stage two: feature extraction 2.1 Large image region feature extraction for each image in : Divide into sub-regions for each sub-region in the upper-level image: Preprocessed sub-region image Extract and store feature vectors 2.2 Global Feature Extraction from Small Images for each small image in : Preprocess small images Extract and store global feature vectors Phase 3: Similarity Matching for each small image in : for each large image in : Calculate the similarity between features of a small image and features of all sub-regions of a large image. Filter out those with similarity higher than the threshold candidate subregions Record the coordinates and similarity of the best-matching sub-region. Phase Four: Verification of Geometric Relationships For each matching relation, input the candidate results: Use feature matching algorithms to verify the geometric consistency of subregions in the small and large images. Mismatches are filtered out using homography matrix calculations. Phase 5: Results Compilation 5.1 Sort matching results by similarity 5.2 Deduplication (Retaining the optimal solution when the same small image corresponds to multiple regions of the same large image) 5.3 Generate standardized subordinate relationship records Phase Six: Output Results 6.1 Returns a set of relationships containing small image index, large image index, region coordinates, and similarity.

[0042] (5) Generate photo roll. By traversing each adjacent set from panoramic to detailed, a complete set of photo relationships from panoramic to detailed images is constructed, thereby realizing the positioning of spatial relationships of different levels of images in the field survey images starting from the panoramic image, and finally forming a photo roll.

[0043] The beneficial effects of the embodiments of the present invention are as follows: 1. High Efficiency and Automation: This invention introduces a fully automated process based on a large model, achieving end-to-end automatic image classification, deduplication, and photo roll generation. Compared to traditional photo roll generation methods, this invention completely eliminates the need for manual operation, greatly improving work efficiency and reducing labor costs.

[0044] 2. Intelligent classification and deduplication: The embodiment of the application uses multi-modal large models and language large models to describe the content and spatial size of different images, and combines image feature extraction networks to deduplicate images in the same level. This not only ensures the accurate classification of images, but also effectively avoids the appearance of duplicate images, ensuring the quality and consistency of the photo album.

[0045] 3. Intelligent spatial relationship analysis: The embodiment of the application can automatically analyze the inclusion relationship between images and automatically generate structured photo albums with spatial dependencies. This intelligent analysis capability makes the generation of photo albums more accurate and better reflects the actual spatial structure of the scene investigation, improving the practicality and reliability of the photo album.

[0046] 4. Improve image resolution adaptability: The embodiment of the application can automatically adjust the image resolution according to actual needs during the generation of the photo album, ensuring that the image resolution in the generated photo album meets the requirements. This solves the problem of inconsistent image resolution caused by manual selection of points and angles in traditional solutions, improving the overall quality of the photo album.

[0047] 5. Reduce manual intervention: The embodiment of the application not only reduces the workload of manual classification, deduplication, and spatial relationship determination of images, but also eliminates the need for manual region selection. This not only reduces the possibility of human error, but also makes the photo album generation process more stable and reliable.

[0048] 6. Adapt to big data environment: In the era of big data, the embodiment of the application can efficiently process large amounts of image data and quickly generate high-quality photo albums. Its automation and intelligence characteristics enable it to handle large-scale image data analysis and processing needs, meeting the requirements of rapid image data analysis in a big data environment.

[0049] Improve user experience: Through automated processes, the embodiment of the application simplifies user operation steps and improves user experience. Users only need to provide original image data, and the rest of the work is automatically completed by the system, greatly reducing the user's burden and improving work efficiency.

[0050] Device embodiment one According to the embodiment of the application, a photo album generation device based on a large model is provided, Figure 4 is a schematic diagram of the photo album generation device based on a large model of the embodiment of the application, as Figure 4 shown, the photo album generation device based on a large model according to the embodiment of the application specifically includes: a title description module 40 for describing the content of all survey images within a predetermined range based on a multi-modal large model according to a first prompt word, generating a detailed title for each survey image; Spatial hierarchy module 42 is used to perform text analysis on the detailed title of each survey image based on the second prompt word according to the large language model, and obtain the spatial hierarchy corresponding to the detailed title. The feature extraction module 44 is used to extract features from the survey images based on the diffusion model network according to the survey images of different spatial levels, and to deduplicate the same view images of the same spatial level. The spatial matching module 46 is used to segment each survey image sequentially, starting from the top spatial level of the survey image set, based on the diffusion model network, to find the most similar survey image in the next spatial level and perform spatial matching. Generation module 48 is used to generate a roll of photos from the complete spatial hierarchy based on the results of spatial matching.

[0051] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operation of each module can be understood with reference to the description of the method embodiments, and will not be repeated here.

[0052] Device Example 2 This invention provides an electronic device, such as... Figure 5 As shown, it includes: a memory 50, a processor 52, and a computer program stored in the memory 50 and executable on the processor 52, wherein the computer program, when executed by the processor 52, performs the steps as described in the method embodiment.

[0053] Device Example 3 This invention provides a computer-readable storage medium storing an information transmission implementation program, which, when executed by a processor 52, performs the steps described in the method embodiment.

[0054] The computer-readable storage media described in this embodiment include, but are not limited to, ROM, RAM, disk, or optical disk.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model-based photo album generation method, characterized by, The method comprises the following steps: Based on the multimodal large model, the content of all survey images within a predetermined range is described according to the first prompt word, and the detail title of each survey image is generated; Based on the large language model, the detail title of each survey image is analyzed according to the second prompt word, and the spatial level corresponding to the detail title is obtained; Based on the diffusion model network, the features of the survey images of different spatial levels are extracted, and the same perspective images of the same spatial level are de-duplicated; Based on the diffusion model network, starting from the survey image set of the uppermost spatial level, each survey image is sequentially segmented, the most similar survey image in the next spatial level is found, and spatial matching is performed; Based on the result of spatial matching, a photo roll from the complete spatial level is generated.

2. The method of claim 1, wherein, Based on the multimodal large model, the content of all survey images within a predetermined range is described according to the first prompt word, and the detail title of each survey image is generated, which specifically comprises: Obtaining a prompt word, wherein the prompt word includes a description of the spatial type of the content shown in the image, a description of the environmental elements of the objects shown in the image, a description of the proportions of the objects that may exist, and a description of the spatial depth of the image; Based on the multi-modal large model, the content of all survey images within a predetermined range is described by using prompt word engineering. Assuming that the set of all survey images of a survey is Then, the image content description of any survey image belonging to the set is ​ Formula 1 ; wherein, is an image under given multi-modal large model prompt , generation temperature control parameter and generation sampling parameter image content description obtained by sampling under the condition.

3. The method of claim 1, wherein, Based on the large language model, the detail title of each survey image is analyzed according to the second prompt word, and the spatial level corresponding to the detail title is obtained, which specifically comprises: Taking the detail title generated for each survey image as input, based on the large language model, according to the given second prompt word, generating temperature control parameters and sampling parameters, judging the spatial level to which each survey image belongs, wherein the second prompt word includes all classification types and the definition of each type, and through examples, the understanding of the large language model for the classification type to be classified is enhanced and the required output format is specified.

4. The method of claim 1, wherein, The spatial level specifically includes: panorama, orientation, overview, and detail.

5. The method of claim 1, wherein, Based on the diffusion model network, the features of the survey images of different spatial levels are extracted, and the same perspective images of the same spatial level are de-duplicated, which specifically comprises: For all survey images in each spatial level, first, the feature extraction network is pre-trained to extract features; on this basis, the similarity between any two features is calculated to construct a similarity matrix between all survey image pairs in the spatial level; based on the preset similarity threshold, the survey images are clustered and grouped to obtain the grouping of all survey images in the image sub-set classified based on the spatial scale according to similarity; after clustering and grouping, the survey image with the highest similarity is selected as the representative survey image under the same perspective in each group.

6. The method of claim 1, wherein, Based on the diffusion model network, starting from the survey image set of the uppermost spatial level, each survey image is sequentially segmented, the most similar survey image in the next spatial level is found, and spatial matching is performed, which specifically comprises: For any two upper and lower level image sets, first, the survey image of the upper level with a larger spatial range is regionally divided into several sub-regions with certain overlapping intervals, and the divided sub-regions contain several different size levels, and after the sub-regions corresponding to each sub-interval are subjected to the cutting preprocessing operation, the pre-trained image feature extraction network is used for feature extraction, and the features and corresponding sub-region attribute information are stored in a dictionary manner; The survey image of the next level with a smaller spatial range is subjected to feature extraction using the image feature extraction network, and for the features corresponding to each survey image of the next level in the spatial range, a sub-region with a similarity exceeding a threshold value in the dictionary of the sub-region features of the previous level image is found, and the attribute information of the best matching sub-region is recorded; after finding the candidate set of the relevant sub-region images in the previous level corresponding to each image of the next level, the geometric consistency is verified by a feature matching algorithm; the results are further de-duplicated to find the spatial relationship between each image in the next level image set and the corresponding large image in the previous level image set, and finally organized into an associated set.

7. The method of claim 1, wherein, Based on the results of spatial matching, a photo album from a complete spatial level is generated, which specifically includes: By traversing each adjacent set from panorama to specific purpose image, a complete photo relationship set from panorama to specific purpose image is constructed, and the positioning of the spatial relationship of different level images in the present survey image starting from the panorama is performed, and finally a photo album is formed. 8.A large model-based photo roll generation apparatus, characterized by, It includes: A title description module for describing the content of all survey images within a predetermined range based on a multi-modal large model according to a first prompt word, generating a detailed title for each survey image; A spatial level module for performing text analysis on the detailed title of each survey image based on a large language model according to a second prompt word to obtain a spatial level corresponding to the detailed title; A feature extraction module for extracting features of survey images based on a diffusion model network according to survey images of different spatial levels, and de-duplicating the same perspective image of the same spatial level; A spatial matching module for starting from the survey image set of the uppermost spatial level, sequentially cutting each survey image, finding the most similar survey image in the next spatial level, and performing spatial matching based on the diffusion model network; A generation module for generating a photo album from a complete spatial level based on the results of spatial matching.

9. An electronic device, comprising: It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, which implements the steps of the large model-based photo album generation method according to any one of claims 1-7 when executed by the processor.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores an information transmission implementation program, and the program is executed by the processor to implement the steps of the large model-based photo album generation method according to any one of claims 1-7.