Information extraction and associated text generation method based on electronic photo

By analyzing electronic photos to obtain information and correlating it with the database, and using AI to generate narrative text, the problem of insufficient relevance of AI graphic and text mutual translation tools in the GIS field is solved, multi-dimensional dynamic narrative display is realized, and user experience and data comprehension capabilities are improved.

CN120234435APending Publication Date: 2025-07-01NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372697.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the field of GIS, the correlation between the content generated by AI graphic and text translation tools and the spatiotemporal database is insufficient, resulting in the lack of scene-based semantics of the output. The multi-source data fusion research fails to effectively integrate user-generated content (UGC), making it impossible to realize image spatiotemporal correlation and multi-dimensional narrative generation.

Method used

By analyzing electronic photos, object information, scene information, date and time information and geographical location information are obtained, and related to the map data in the database, narrative text is generated using AI models to realize "position-time-scene" three-dimensional dynamic narrative, and visually display it through GIS applications.

Benefits of technology

Generate scene-based text, increase storytelling, improve user experience, realize dynamic visualization of static archives, and expand applications in fields such as culture and education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234435A_ABST
    Figure CN120234435A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of geographic information systems (GIS), artificial intelligence (AI) and cross-modal data fusion, in particular to an information extraction and associated text generation method based on electronic photos. According to the method, an electronic photo set is analyzed and processed in sequence, and multi-dimensional information including object information, scene information, date and time information and geographic position information is obtained through various methods such as an image analysis technology, timestamp extraction and space coordinate analysis. The electronic photo is further associated with map data in a database, including but not limited to geographical name and historical geographical name information. And finally, inputting the object information, the scene information, the date and time information, the geographic position information, the place name information and the historical place name information into an AI model as cue words, intelligently generating a narrative text, realizing'position-time-scene 'three-dimensional dynamic narrative, and visually displaying a position association jump mode through a GIS application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of geographic information systems (GIS), artificial intelligence (AI), and cross-modal data fusion, and particularly relates to a method for information extraction and associated text generation based on electronic photos. Technical Background

[0002] In recent years, the technology of artificial intelligence and multi-source spatio-temporal data fusion has shown significant application potential in the field of geographic information systems (GIS). [1-2] With the breakthrough progress of deep learning technology, AI image recognition technology has made important progress in the directions of object detection and scene understanding. [3] Research shows that data augmentation methods based on OpenCV can significantly improve the object recognition accuracy in complex scenes. [4] The introduction of geometric invariance features provides a new theoretical framework for three-dimensional object recognition. [5] Meanwhile, the technology for extracting image Exif information has become mature, enabling the parsing and extraction of image metadata (latitude and longitude coordinates, timestamp). [6]

[0003] Although the existing technology already has the ability to process multi-modal data, key bottlenecks still exist: First, the generated content of mainstream AI image-text translation tools has insufficient association with the spatio-temporal database, resulting in the output lacking scene-based semantics. [7] Second, multi-source data fusion research mostly focuses on structured data analysis and rarely involves the dynamic integration of user-generated content (UGC). For example, the multi-source spatio-temporal data aggregation method proposed by Interstellar Space (Tianjin) Technology Development Co., Ltd. has achieved the conversion of non-spatial data to spatial data, but has not solved the problems of image spatio-temporal association and multi-dimensional narrative generation. [8] The traffic flow prediction model of the Guangzhou Urban Planning and Surveying and Design Institute has improved the prediction accuracy through spatio-temporal correlation fusion, but has not integrated text generation technology. [9]

[0004] To improve the existing situation, the present invention proposes a method for information extraction and associated text generation based on electronic photos. By analyzing electronic photos, object information, scene information, date and time information, and geographical location information of the electronic photos are obtained. By associating the electronic photos with map data in the database, including but not limited to place names and historical place name information. The above information is input into a cross-modal AI model as prompt words to achieve three-dimensional dynamic narrative of "location-time-scene", and the location association jump method is visually displayed through GIS applications. The present invention can not only break through the static information display of traditional maps, but also be applied in multiple scenarios such as cultural protection and environmental change, providing a new technical paradigm for smart city and digital humanities research.

[0005] References

[0006] [1] Jigen Lin, Bin Zhao. Review of Spatiotemporal Data Mining for Big Data [J]. Journal of Nanjing Normal University (Natural Science Edition), 2014, 37(01): 1 - 7.

[0007] [2] Binbin Zhao, Guangqiang Li, Min Deng. Review of Spatiotemporal Data Mining [J]. Science of Surveying and Mapping, 2010, 35(02): 62 - 65 + 20. DOI: 10.16251 / j.cnki.1009 - 2307.2010.02.058.

[0008] [3] Deshan Yu. Visual Truth and Its Reflection in the Era of Artificial Intelligence [J]. Social Science Front, 2020, (01): 224 - 233.

[0009] [4] Xiaokai Sun, Qingyuan Ni, Wenqiang Chen. Feasibility Study on the Application of Image Enhancement Methods in Deep Learning Image Recognition Scenarios [J]. Telecommunications Science, 2020, 36(S1): 172 - 179.

[0010] [5] Kaiqi Huang, Weiqiang Ren, Tieniu Tan. Review of Image Object Classification and Detection Algorithms [J]. Chinese Journal of Computers, 2014, 37(06): 1225 - 1240.

[0011] [6] Chuan Shen. Analysis of the Application of Exif Information in the Authenticity Identification of Digital Photos [J]. Imaging Technology, 2017, 29(03): 74 - 76.

[0012] [7] Baiyang Li, Yun Bai, Xini Zhan, etc. Technical Characteristics and Morphological Evolution of Artificial Intelligence - Generated Content (AIGC) [J]. Document, Information & Knowledge, 2023, 40(01): 66 - 74. DOI: 10.13366 / j.dik.2023.01.066.

[0013] [8] Wenbo Yang, Aohui Ma, Hongyue Liu, etc. A Method for Converging and Sharing Multi - source Spatiotemporal Data [P]. Tianjin: CN202311448569.X, 2025 - 01 - 14.

[0014] [9] Xingdong Deng, Guanyao Li, Yang Liu, etc. Traffic Flow Prediction Method, Device and Storage Medium Integrating Multi - source Spatiotemporal Data [P]. Guangdong: CN202310084450.2, 2025 - 01 - 17. Summary of the Invention

[0015] The present invention provides a method for information extraction and associated text generation based on electronic photos, characterized in that the electronic photo set is sequentially analyzed and processed to obtain object information, scene information, date and time information, and geographical location information of the electronic photos. By associating the electronic photos with map data in the database, including but not limited to place names and historical place name information. The object information, scene information, date and time information, geographical location information, and place name and historical place name information are input into the AI model as prompts to intelligently generate narrative text, realizing three-dimensional dynamic narrative of "location-time-scene", and visually displaying the location association jump method through GIS application.

[0016] Specifically, it includes:

[0017] The electronic photo set (A) contains the entire sequence of electronic photos to be processed, and the electronic pictures contain geographical coordinate information;

[0018] Processing the electronic photos includes image analysis (a), using image recognition technology to extract object information and scene information from the electronic photos;

[0019] Processing the electronic photos includes date and time parsing (b), using a computer program to automatically obtain the electronic photo tag information and further extract the date and time information;

[0020] Processing the electronic photos includes latitude and longitude coordinate parsing (c), using a computer program to automatically obtain the electronic photo tag information and further extract the latitude and longitude coordinate information;

[0021] AI generation (d) is to use the AI model to intelligently generate narrative text according to the input prompts in chronological order;

[0022] The database (B) is used to store geospatial data and metadata;

[0023] The database (B) is also used to establish an association relationship with geographical location information, including place names and historical place name information;

[0024] The client (C) integrates an online map (1) and a graphic story (2);

[0025] The online map (1) is used to realize the annotation display of geographical location information on the map and support map layer switching;

[0026] The graphic story (2) is used to display the generated narrative text and support user interaction.

[0027] The specific steps of the present invention include:

[0028] S1 Image analysis (a): Obtain object information and scene information of the electronic photos through image analysis technology;

[0029] S2 Date and Time Parsing (b): Automatically obtain the date and time information of electronic photos through a computer program;

[0030] S3 Coordinate Parsing (c): Automatically obtain the GPS longitude and latitude coordinate information of electronic photos through a computer program;

[0031] S4 Geographic Information Association: Match and associate electronic photos with map data in database (B) through longitude and latitude coordinate information, mark coordinate points on the map, and establish an association relationship between the marked points and geographic location information, including place names and historical place name information;

[0032] S5 Text Generation: Use the object information, scene information, date and time information, geographic location information, place names, and historical place name information of electronic photos as prompts, input them into the AI model, and generate (d) narrative text by the AI;

[0033] S6 Visualization Display: The client set (C) integrates GIS applications and content display, receives the map annotation results and the narrative text generated (d) by the AI in real time, and supports user interaction operations.

[0034] Preferably, an information extraction and associated text generation method based on electronic photos is characterized in that:

[0035] This method uses a computer program to achieve automated processing and outputs the processing results in real time.

[0036] Preferably, an information extraction and associated text generation method based on electronic photos is characterized in that:

[0037] The narrative text generated and output by the AI text contains the following key information:

[0038] Geographic location association, including place names and historical place names;

[0039] Time acquisition, the date and time information determines the narrative order of the electronic photo set;

[0040] Scene content interpretation, semantic text description based on object information and scene information;

[0041] Historical information association, description of the source of historical place names and historical stories.

[0042] Preferably, for step S4 geographic location information association, including place name and historical place name information, it is characterized in that:

[0043] According to the geospatial data matched with the longitude and latitude coordinate information in the database, obtain the attribute elements, and further associate the geographic location information, including place name and historical place name information.

[0044] Preferably, in the S6 visual display, the client integrates GIS applications and content display, characterized in that:

[0045] Clicking on the electronic photo or place name text triggers an event, and the map view moves to the map marked point;

[0046] Switch between historical / modern map layers for comparative display.

[0047] The beneficial effects of the present invention are as follows:

[0048] Associating spatial location, real-time environment, and historical events through a unified model to generate scenario-based text, enhancing the storytelling and improving the user experience.

[0049] Generate content with three-dimensional dynamic narrative of "location-time-scene" and temporality, helping users better understand and describe how data or events change over time and space.

[0050] Dynamic visualization of historical data: Through AI and visualization technologies, convert static archives into interactive content, expand the spatio-temporal switching of scenarios, and can be applied to fields such as culture and education. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.

[0052] In the drawings:

[0053] Figure 1 is a flowchart of an embodiment of the present invention.

[0054] In the figure: electronic photo collection (A), database (B), client (C), image analysis (a), date and time parsing (b), coordinate parsing (c), AI generation (d), online map (1), graphic story (2). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0056] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] Embodiment:

[0058] Please refer to Figure 1, the present invention provides a method for information extraction and associated text generation based on electronic photos. By parsing and processing electronic photos, object information, scene information, date and time information, and geographical location information are obtained. An associated relationship between the electronic photos and map data in the database, including place names and historical place name information, is established. And all the above information is input into an AI model as prompt words to intelligently generate narrative text, and the position association jump method is visually displayed through GIS application.

[0059] The present invention is further described below with reference to the accompanying drawings:

[0060] The electronic photo set (A) contains the entire sequence of electronic photos to be processed, and the electronic pictures contain geographical coordinate information;

[0061] Processing the electronic photos includes image analysis (a), using image recognition technology to extract object information and scene information in the electronic photos;

[0062] Processing the electronic photos includes date and time parsing (b), using a computer program to automatically obtain the electronic photo tag information and further extract the date and time information;

[0063] Processing the electronic photos includes latitude and longitude coordinate parsing (c), using a computer program to automatically obtain the electronic photo tag information and further extract the latitude and longitude coordinate information;

[0064] AI generation (d) is to use an AI model to intelligently generate narrative text in chronological order according to the input prompt words;

[0065] The database (B) is used to store geospatial data and metadata;

[0066] The database (B) is also used to establish an associated relationship with geographical location information, including place names and historical place name information;

[0067] The client (C) integrates an online map (1) and a graphic story (2);

[0068] The online map (1) is used to realize the annotation display of geographical location information on the map and support map layer switching;

[0069] The graphic story (2) is used to display the generated narrative text and support user interaction.

[0070] The specific steps of the present invention include:

[0071] S1 Image analysis (a): Obtain object information and scene information of electronic photos through image analysis technology;

[0072] S2 Date and time parsing (b): Automatically obtain the date and time information of electronic photos through a computer program;

[0073] S3 Coordinate Analysis (c): Automatically obtain the GPS latitude and longitude coordinate information of the electronic photo through a computer program;

[0074] S4 Geographic Information Association: Match and associate the electronic photo with the map data in the database (B) through the latitude and longitude coordinate information, mark the coordinate points on the map, and establish the association relationship between the marked points and the geographic location information, including place names and historical place name information;

[0075] S5 Text Generation: Use the object information, scene information, date and time information, geographic location information, place names and historical place name information of the electronic photo as prompt words, input them into the AI model, and generate (d) narrative text by the AI

[0076] S6 Visualization Display: The client (C) integrates GIS applications and content display, receives the map marking results and the narrative text generated (d) by the AI in real time, and supports user interaction operations.

[0077] In the example:

[0078] Each electronic photo to be processed in the electronic photo set (A) has a unique ID;

[0079] The unique ID is used to ensure that each electronic photo has an independent identifier for distinction;

[0080] The database (B) includes a basic map database and a historical spatio-temporal database;

[0081] The basic geographic database includes a high-precision vector base map, a topographic base map, a topographic shaded relief, and a satellite map;

[0082] The historical spatio-temporal database includes historical map versions and electronic photo metadata;

[0083] The seamless superposition of multi-version map data is realized through a spatio-temporal alignment algorithm.

[0084] The S1 Image Analysis (a) mainly includes:

[0085] Object Detection: Use an object detection algorithm to identify and locate the position and bounding box of the object in the electronic photo. For each detected object, determine its position, size, and shape information;

[0086] Object Recognition: Based on the detected object information, use a deep learning model to classify and recognize each object to determine the specific category of the object;

[0087] Scene Analysis: Use a scene analysis algorithm to identify the scene of the entire electronic photo and judge the background and environmental types in the electronic photo;

[0088] Feature extraction: For each detected object and recognized scene, extract the key feature description information, that is, object information and scene information.

[0089] The S2 date and time parsing (b) mainly includes:

[0090] Automatically obtain the tag information of the electronic photo using a computer program;

[0091] Search for the date and time tags and parse the date and time information of the taken electronic photo.

[0092] The S3 coordinate parsing (c) mainly includes:

[0093] Automatically obtain the tag information of the electronic photo using a computer program;

[0094] Parse the GPS latitude and longitude coordinate information in the electronic photo and perform corresponding conversion processing to convert the degree-minute-second representation to the decimal representation;

[0095] Judge the reference direction of the latitude and longitude (north latitude / south latitude, east longitude / west longitude) and perform positive and negative sign adjustment;

[0096] Output the obtained latitude and longitude coordinate information.

[0097] The S4 geographic information association mainly includes:

[0098] Map annotation: Match the obtained latitude and longitude coordinate information with the geospatial data in the database (B), automatically create annotation points using a computer program, and add the annotation points to the map for display through the map service interface;

[0099] Electronic photo association: Establish an association relationship between the annotation points and the electronic photos according to the annotation points;

[0100] Place name association: Overlay the vector map in the database (B) according to the annotation points, and match and associate the place name information of the corresponding geographical location;

[0101] Historical place name association: Overlay the historical map version in the database (B) according to the annotation points, and match and associate the historical place name information of the corresponding geographical location;

[0102] Geographical location information, including place names and historical place name information.

[0103] The S5 text generation mainly includes:

[0104] Input the object information, scene information, date and time information, geographical location information, place names and historical place name information as prompt words into the AI model;

[0105] Based on the input prompt words, the AI generates relevant narrative text (d) to describe information such as objects, scenes, dates and times, geographical locations, etc. in the digital photo;

[0106] The AI model can sequentially present the generated content according to the timeline, describe the content changes of geographical location points over time, and form a dynamic narrative.

[0107] The S6 visualization display mainly includes:

[0108] The client (C) integrates an online map (1) and a graphic story (2);

[0109] The online map (1) realizes the display of marked points on the map through an online map service interface;

[0110] The online map (1) also supports the function of switching map layers so that users can view the marked points under different map types;

[0111] The graphic story (2) automatically realizes the display of the narrative text generated by the AI through a computer program, and supports interactive operations such as user browsing, commenting, and sharing;

[0112] According to the association relationship between the digital photo and the marked point established in S4, the client (C) supports the interactive operation of moving the map view to the specified marked point by clicking on the digital photo;

[0113] According to the association relationship between the place name text and the marked point established in S4, the client (C) supports the interactive operation of moving the map view to the specified marked point by clicking on the place name text;

[0114] According to the association relationship between the historical place name text and the marked point established in S4, the client (C) supports the interactive operation of moving the map view to the specified marked point by clicking on the historical place name text.

[0115] The above shows and describes the basic principles, main features, and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0116] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for extracting information and generating associated text based on electronic photos, characterized in that: The electronic photo collection is analyzed and processed in sequence to obtain the object information, scene information, date and time information and geographic location information of the electronic photos. By associating the electronic photos with the map data in the database, including but not limited to place names and historical place name information. The object information, scene information, date and time information, geographic location information and place names and historical place name information are input into the AI ​​model as prompt words, and narrative text is intelligently generated to achieve a three-dimensional dynamic narrative of "location-time-scene", and the location association jump method is visualized through GIS application.

2. A method for extracting information and generating associated text based on electronic photos according to claim 1, comprising the following steps: S1 Image Analysis: Obtain object information and scene information of electronic photos through image analysis technology; S2 Date and time analysis: Automatically obtain the date and time information of electronic photos through computer programs; S3 coordinate analysis: Automatically obtain the GPS longitude and latitude coordinate information of electronic photos through computer programs; S4 Geographic Information Association: Match and associate electronic photos with map data in the database through longitude and latitude coordinate information, mark coordinate points on the map, and establish an association between the marked points and geographic location information, including place names and historical place name information; S5 text generation: The object information, scene information, date and time information, geographic location information, place names and historical place name information of the electronic photo are used as prompt words and input into the AI ​​model, and the AI ​​generates narrative text; S6 Visualization Display: The client integrates GIS applications and content display, receives map annotation results and AI-generated narrative text in real time, and supports user interactive operations.

3. A method for extracting information and generating associated text based on electronic photos according to claim 1, characterized in that: The method utilizes a computer program to realize automatic processing and outputs the processing results in real time.

4. A method for extracting information and generating associated text based on electronic photos according to claim 1, characterized in that: The narrative text generated by the AI ​​text generation output contains the following key information: Geographical associations, including place names and historical place names; Time acquisition, where date and time information determines the narrative order of the electronic photo collection; Scene content interpretation, based on semantic text description of object information and scene information; Association of historical information, origin of historical place names and description of historical stories.

5. According to step S4 of claim 2, the geographical location information is associated with the place name and the historical place name information, characterized in that: Based on the geospatial data that matches the longitude and latitude coordinate information in the database, attribute elements are obtained to further associate geographic location information, including place names and historical place name information.

6. According to step S6 of claim 2, the client integrates GIS application and content display, characterized in that: Click on the electronic photo or place name text to trigger the event, and the map view moves to the map marked point; Toggle historical / modern map layer comparison display.