Method for generating training data for generative artificial intelligence using spatial information and real-time sensor information and apparatus for same
By converting location-based object data into documented data and using it as learning input for generative AI, the method addresses the limitations of existing AI models in handling location-related and real-time queries, significantly improving response accuracy.
Patent Information
- Application Number
- PCT/KR2024/096074
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-12
AI Technical Summary
Generative artificial intelligence models struggle to accurately respond to location-related user queries and real-time data that has not been pre-trained, particularly in environments where data is not documented in text form.
A method for generating learning data for generative AI by converting location-based object data, such as maps, sensors, weather, and performances, into documented data in real-time using spatial information and real-time sensor data, and then providing this data as learning input for AI models.
This approach enhances the accuracy of generative AI responses to user queries by incorporating real-time and location-specific data, improving the model's ability to handle queries that were previously unanswerable or inaccurately answered.
Smart Images

Figure KR2024096074_12062025_PF_FP_ABST
Abstract
Description
A method and device for generating learning data for generative artificial intelligence using spatial information and real-time sensor information.
[0001] The present invention relates to a method and a device for generating learning data for generative artificial intelligence using spatial information and real-time sensor information, and relates to a method for converting data that does not exist in text form into documented data in text form and providing it as learning data for generative AI.
[0002] The material described in this section merely provides background information on one embodiment of the present invention and does not constitute prior art.
[0003] Large language models (LLMs) and other large AIs (artificial intelligence) are artificial intelligence models trained using a very large amount of text data, perform natural language processing tasks, and can be used for various language modeling tasks.
[0004] These super-large AIs are the largest and most complex models in the field of artificial intelligence, and they demonstrate very high performance in various fields of artificial intelligence, such as natural language processing, image recognition, and speech recognition.
[0005] A representative example of a large-scale AI is the GPT (generative pretrained transformer) series developed by Open AI, which has shown outstanding performance in the field of natural language processing and has a high level of language understanding and generation ability through learning using large amounts of text data.
[0006] Furthermore, super-large-scale AI is already being used in a variety of fields. For example, in natural language processing, it can be applied to various applications such as automatic translation, automatic summarization, and question answering. In image analysis, it can be applied to image classification and object recognition. In speech recognition, it can be used for speech synthesis and voice recognition.
[0007] This development of super-large AI is largely due to advances in deep learning algorithms and hardware technology. Specifically, advances in deep learning algorithms have enabled super-large AI models to learn more complex patterns, while advances in hardware technology have significantly accelerated the learning speed of these models.
[0008] LLM can be trained on a training dataset consisting of hundreds of billions of sentences from various web documents and text data, such as the Internet, books, newspaper articles, and blogs. It can be used in various applications such as natural language understanding, sentence generation, machine translation, chatbots, and automatic summarization. These LLM-based deep learning models are artificial intelligence technologies that learn human language and perform various tasks using that language.
[0009] Deep learning models based on LLM are representative of tasks that understand human-written sentences through natural language understanding and answer questions based on this understanding. They are particularly attracting great attention in fields such as chatbots, artificial intelligence that converses with humans.
[0010] Representative LLM-based deep learning models, such as OpenAI's ChatGPT and Google's Bard, are trained on internet documents. Therefore, their performance is bound to be poor when responding to real-time data generated in situations where prior training has not been conducted. In particular, LLM-based deep learning models have a significant problem: their accuracy in responding to location-related user queries is significantly reduced, especially for small and medium-sized businesses and businesses that experience approximately one million business openings and closures annually.
[0011] Thus, for user queries related to data that is not documented in text form or data generated in real time for monitoring, deep learning models based on LLM have the problem of being unable to answer or of producing significantly low accuracy results.
[0012] In particular, generative AI used in conversational AI services has the problem that learning itself cannot be performed in the case of map data that does not exist in text form on the Internet and real-time data occurring at the present time due to the limitations of language models learned based on data before the user inputs a question, and as a result, most questions related to location information or environmental information closely related to our lives are either impossible to answer or present answers with many errors.
[0013] Prior art literature includes Korean Patent No. 10-2570178 (August 21, 2023).
[0014] The present invention has been devised in response to the aforementioned background technology, and aims to provide a method for converting location-based object data, such as maps, sensors, weather, and performances, which are not textualized on the Internet, such as the Internet and social networks, into documented data in real time and providing the data as learning data for a generative artificial intelligence model.
[0015] However, the problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned can be clearly understood based on the description below.
[0016] As a technical means for achieving the above-described technical task, an object of the present invention is to provide a method for generating learning data of a generative artificial intelligence, which is performed by a computing device including at least one processor according to an embodiment of the present invention. The method comprises the steps of: obtaining location-based object data for an object including a preset object; converting metadata and object attribute data specifying the object based on the location-based object data into a textual listing format to generate documented data; and generating learning data based on the documented data.
[0017] Alternatively, the step of converting the metadata and object property data that specify the object based on the location-based object data into a text listing format to create documented data is, when the object includes at least one unit object, to create a string that converts the metadata and unit object property data that specify the unit object into a text listing format for each unit object, and merges the strings for each unit object to create documented data.
[0018] Alternatively, the location-based object data is at least one of spatial data based on any one of points, polygons or polylines, topographic data based on digital elevation model (DEM) data, sensor data generated through real-time environmental monitoring, public data using an open application programming interface (API), or spatial analysis data based on three-dimensional spatial information.
[0019] Alternatively, the metadata includes at least one piece of information from among coordinate information of a unit object, address information structured by geocoding the coordinate information, absolute altitude information of the coordinate information, and location characteristic information including an average slope, slope direction, or distance to surrounding roads or major facilities for a predetermined change area based on the coordinate information, when the location-based object data includes spatial data based on the point.
[0020] Alternatively, the step of converting the metadata and object attribute data that specify the object based on the location-based object data into a text listing format to create documented data is to generate a document string for each unit object by separating the metadata converted into a text listing format and the unit object attribute data with a delimiter.
[0021] Alternatively, the metadata includes at least one piece of information from among address information structured by geocoding based on the central coordinate information of the polygon, and location characteristic information including the area, slope, slope direction, or elevation above sea level of the polygon, when the location-based object data includes spatial data based on the polygon.
[0022] Alternatively, the metadata includes at least one piece of information from among coordinate information including the latitude and longitude coordinates of the center point of the polyline and the coordinates of the start and end points, address information structured by geocoding the coordinate information, and location characteristic information including the line length of the polyline, the line slope or the average altitude of the line, or surrounding environment information located in a preset change area based on the coordinate information, when the object data includes spatial data based on the polyline.
[0023] Alternatively, the metadata includes at least one of information among: center coordinate information of a coordinate group defined based on the unit column when the object data includes the terrain data, address information structured by geocoding based on the center coordinate information, and location characteristic information including elevation, slope, slope direction, or land cover for a change area set based on the center coordinate information.
[0024] Alternatively, the metadata includes location information and measurement values of sensors collected at preset time units when the location-based object data includes the sensor data.
[0025] Alternatively, the sensor is an IoT sensor including at least one of a fine dust sensor, a bridge vibration monitoring sensor, or a traffic volume measurement monitoring sensor.
[0026] Alternatively, the step of converting metadata and object attribute data specifying the object based on the location-based object data into a text listing format and generating documented data includes, when the location-based object data includes public data, the step of generating basic learning data based on at least one public API; the step of learning a first language model based on the basic learning data; the step of determining a public API matching a user query using a pre-learned first language model, and then calling the determined public API to receive a result value; and the step of converting the result value into documented data and providing it as a prompt of the first language model.
[0027] Alternatively, the metadata may include at least one of spatial extent information including layers and lakes for the object, or visibility information regarding the ratio of one or more environmental elements based on the visibility analysis results of the object, if the object data includes the spatial analysis data.
[0028] Meanwhile, as a technical means for achieving the above-described technical task, an object of the present invention is to provide a computing device for generating learning data of a generative artificial intelligence according to an embodiment of the present invention. The device comprises: a processor including at least one core; and a memory including program codes executable by the processor; wherein the processor, upon execution of the program code, obtains location-based object data for an object including a preset object, converts metadata and object attribute data specifying the object based on the location-based object data into a textual listing format to generate documented data, and generates learning data based on the documented data.
[0029] According to the problem solving means of the present invention described above, the present invention can convert location-based object data such as maps, sensors, weather, and performances that are not textualized on the Internet, such as the Internet and social networks, into documented data in real time and provide it as learning data for a generative artificial intelligence model, and since the generative artificial intelligence model is retrained based on the documented data, the generative artificial intelligence model can improve the accuracy of answers to user queries.
[0030] FIG. 1 is a block diagram of a computing device according to one embodiment of the present invention.
[0031] FIG. 2 is a flowchart illustrating a method for generating learning data of a generative artificial intelligence according to one embodiment of the present disclosure.
[0032] FIG. 3 is a diagram illustrating a learning data generation process for spatial data based on points according to one embodiment of the present disclosure.
[0033] FIG. 4 is a diagram illustrating a learning data generation process for spatial data based on polygons according to one embodiment of the present disclosure.
[0034] FIG. 5 is a diagram illustrating a learning data generation process for spatial data based on a polyline according to one embodiment of the present disclosure.
[0035] FIG. 6 is a diagram illustrating a learning data generation process for terrain data according to one embodiment of the present disclosure.
[0036] FIG. 7 is a diagram illustrating a learning data generation process for sensor data according to one embodiment of the present disclosure.
[0037] FIG. 8 is a diagram illustrating a learning data generation process for public data according to one embodiment of the present disclosure.
[0038] Below, embodiments of the present invention are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present invention. The embodiments presented in the present invention are provided to enable those skilled in the art to utilize or practice the contents of the present invention. Accordingly, various modifications to the embodiments of the present invention will be apparent to those skilled in the art. That is, the present invention can be implemented in various different forms and is not limited to the embodiments described below.
[0039] Throughout the specification of the present invention, identical or similar drawing numbers refer to identical or similar components. Furthermore, for the purpose of clearly explaining the present invention, drawing numbers for parts in the drawings that are not relevant to the description of the present invention may be omitted.
[0040] The term "or" as used herein is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified herein or the meaning is clear from the context, "X utilizes A or B" should be understood to mean either of the natural inclusive permutations. For example, unless otherwise specified herein or the meaning is clear from the context, "X utilizes A or B" can be interpreted to mean either X utilizes A, X utilizes B, or X utilizes both A and B.
[0041] The term “at least one of A or B” as used in the present invention should be interpreted to refer to all of A, B, and combinations of A and B.
[0042] The term "and / or" as used herein should be understood to refer to and include all possible combinations of one or more of the related concepts listed.
[0043] The terms "comprises" and / or "comprising" as used herein should be understood to mean the presence of certain features and / or components. However, it should be understood that the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other features, other components, and / or combinations thereof.
[0044] Unless otherwise specified in the present invention or unless the context makes it clear that the singular form is indicated, the singular should generally be construed to include "one or more."
[0045] The term "Nth (N is a natural number)" used in the present invention can be understood as an expression used to mutually distinguish components of the present invention based on a predetermined standard such as a functional perspective, a structural perspective, or convenience of explanation. For example, components performing different functional roles in the present invention can be distinguished as a first component or a second component. However, components that are substantially the same within the technical spirit of the present invention but must be distinguished for convenience of explanation may also be distinguished as a first component or a second component.
[0046] Meanwhile, the term "module" or "unit" used in the present invention can be understood as a term referring to an independent functional unit that processes computing resources, such as a computer-related entity, firmware, software or a part thereof, hardware or a part thereof, or a combination of software and hardware. At this time, the "module" or "unit" may be a unit composed of a single element, or a unit expressed as a combination or set of multiple elements. For example, as a narrow concept, a "module" or "unit" may refer to a hardware element of a computing device or a set thereof, an application program that performs a specific function of software, a processing process implemented through software execution, or a set of instructions for program execution, etc. In addition, as a broad concept, a "module" or "unit" may refer to the computing device itself that constitutes the system, or an application that runs on the computing device, etc. However, since the above-described concept is only an example, the concept of “module” or “part” can be defined in various ways within a range understandable to those skilled in the art based on the contents of the present invention.
[0047] The explanation of the aforementioned terms is intended to aid understanding of the present invention. Therefore, unless explicitly stated as limiting the scope of the present invention, it should be noted that the aforementioned terms are not intended to limit the technical concept of the present invention.
[0048]
[0049] FIG. 1 is a block diagram of a computing device according to one embodiment of the present invention.
[0050] The computing device (100) according to one embodiment of the present invention may be a hardware device or a part of a hardware device that performs comprehensive processing and calculation of data, or may be a software-based computing environment connected to a communication network. For example, the computing device (100) may be a server that performs intensive data processing functions and shares resources, or may be a client that shares resources through interaction with a server. In addition, the computing device (100) may be a cloud system that enables multiple servers and clients to interact with each other to comprehensively process data. Since the above description is only one example related to the type of the computing device (100), the type of the computing device (100) may be configured in various ways within a range understandable to those skilled in the art based on the contents of the present invention.
[0051] Referring to FIG. 1, a computing device (100) according to one embodiment of the present invention may include a processor (110), a memory (120), and a network unit (130). However, FIG. 1 is merely an example, and the computing device (100) may include other components for implementing a computing environment. In addition, only some of the disclosed components may be included in the computing device (100).
[0052] The processor (110) according to one embodiment of the present invention may be understood as a configuration unit including hardware and / or software for performing computing operations. For example, the processor (110) may read a computer program and perform data processing through computational functions and control functions. The processor (110) for performing such data processing may include a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). The type of the processor (110) described above is only one example, and thus, the type of the processor (110) may be configured in various ways within a range understandable to those skilled in the art based on the contents of the present invention.
[0053] This processor (100) can generate response data for user queries regarding location-based object data using a pre-trained first language model. Furthermore, the processor (110) can retrain the first language model and the second language model based on the response data using the first language model.
[0054] In this case, the first language model may be a generative artificial intelligence (AI) model, and the second language model may be a large language model (LLM). Accordingly, the second language model may be used in the first language model to generate content based on human language input prompts. Accordingly, the processor (110) may use the first language model based on the second language model to generate an agent that responds to user queries.
[0055] Here, the location-based object data can be any of spatial data based on points, polygons, or polylines, topographic data based on digital elevation model (DEM) data, sensor data for real-time environmental monitoring, public data using an open application programming interface (API), or spatial analysis data based on three-dimensional spatial information.
[0056] Location-based object data can include coordinates, text information such as altitude and slope among spatial information, and non-text information that is not textualized.
[0057] The processor (110) can perform an operation to express at least one neural network block included in the first language model and the second language model during the learning process of the neural network model.
[0058] The processor (110) can extract text information to be used as training data for the first language model using the second language model generated through the above-described learning process, and can generate training data to be used for the first language model based on the extracted text information. The processor (110) can input a user query into the first language model trained through the above-described process, and generate response data representing an estimated result reflecting knowledge information at the time of the query.
[0059] In addition to the examples described above, the types of learning datasets according to language distribution and the outputs of the first language model and the second language model can be configured in various ways within a range understandable to those skilled in the art based on the contents of the present disclosure.
[0060] The memory (120) according to one embodiment of the present invention may be understood as a configuration unit including hardware and / or software for storing and managing data processed in the computing device (100). That is, the memory (120) may store any type of data generated or determined by the processor (110) and any type of data received by the network unit (130). For example, the memory (120) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory, a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, or an optical disk. In addition, the memory (120) may also include a database system that controls and manages data in a predetermined system. The type of memory (120) described above is only one example, and thus the type of memory (120) can be configured in various ways within a range understandable to those skilled in the art based on the contents of the present invention.
[0061] The network unit (130) according to one embodiment of the present invention can be understood as a component that transmits and receives data through any type of known wired or wireless communication system. For example, the network unit (130) can perform data transmission and reception using a wired or wireless communication system such as a local area network (LAN), wideband code division multiple access (WCDMA), long term evolution (LTE), wireless broadband internet (WiBro), fifth generation mobile communication (5G), ultra-wideband, ZigBee, radio frequency (RF) communication, wireless LAN, wireless fidelity, near field communication (NFC), or Bluetooth. Since the above-described communication systems are only examples, the wired and wireless communication system for data transmission and reception of the network unit (130) can be applied in various ways other than the above-described examples.
[0062]
[0063] FIG. 2 is a flowchart illustrating a method for generating learning data of a generative artificial intelligence according to one embodiment of the present disclosure.
[0064] Referring to FIG. 2, a computing device (100) obtains location-based object data including text information or non-text information about an object including a preset object (S100).
[0065] The computing device (100) converts metadata and object attribute data that identify an object based on the acquired location-based object data into a textual listing format to generate documented data (S200). At this time, the computing device (100) can convert unstructured text into a structured format using a pre-trained second language model.
[0066] When an object includes at least one unit object, the computing device (100) generates a string by converting metadata and unit object attribute data that specify each unit object into a text listing format, and merges the strings for each unit object to generate documented data.
[0067] At this time, the computing device (100) generates a document string for each object by separating the metadata and object attribute data converted into a text listing format with a delimiter such as a comma (,) or special symbol (*, #).
[0068] The computing device (100) generates learning data based on documented data and provides the generated learning data to the first language model, thereby enabling the first language model to be trained based on the learning data (S300).
[0069]
[0070] FIG. 3 is a diagram illustrating a learning data generation process for spatial data based on points according to one embodiment of the present disclosure.
[0071] Vectors and rasters are data types used to represent spatial data. Vector time represents the real world in geometric forms, such as points, polylines, and polygons, while raster represents the real world in a grid of pixels (or cells). Points are composed of latitude and longitude coordinates in most coordinate systems and are useful when the data is too small to be represented as a polygon. For example, points can be used to represent the location of Seoul on a world map. Polylines consist of two or more vertices connected by edges and are used to represent natural linear shapes. For example, a polyline can be used to represent subway lines on a map of Seoul. Polygons consist of three or more vertices in a closed, connected form and are used to represent the boundaries of a specific area. For example, a polygon can be used to mark the boundaries of Seoul on a map of South Korea.
[0072] In the case of point-based spatial data, the metadata may include one or more of the following information: coordinate (x, y) information of the unit object, structured address information by geocoding the coordinate information, absolute elevation information of the coordinate information, and location characteristic information including average slope, slope direction, and distance to surrounding roads or major facilities for a predetermined change area based on the coordinate information.
[0073] Here, the unit object may be a unit column reference point of the learning data, and the major facilities may include convenience facilities such as schools, subway stations, bus stops, educational facilities, medical facilities, restaurants, financial institutions, etc.
[0074] Accordingly, the computing device (100) extracts a unit column reference point (S211), inputs metadata in text form, and generates a string for each unit object by separating the object attribute data previously held by the point information with a delimiter (S212).
[0075] The computing device (100) inputs metadata and object attribute data for all unit objects in point-based spatial data to generate a string for each unit object, and then merges the generated strings for each unit object to generate documented data (S213).
[0076] The first language model can be trained using training data based on documented data about points. Therefore, when a user query such as "Show me a list of Chinese restaurants within 10 minutes of Gangnam Station" is entered from a user terminal or application, the first language model can provide response data to the user query through AI-based metadata retrieval.
[0077]
[0078] FIG. 4 is a diagram illustrating a learning data generation process for spatial data based on polygons according to one embodiment of the present disclosure.
[0079] As illustrated in FIG. 4, in the case of spatial data based on a polygon, the computing device (100) inputs metadata including at least one piece of information from among structured address information, area of the polygon, slope, slope direction, and location characteristic information including the altitude above sea level, by geocoding based on the central coordinate information of the polygon.
[0080] The computing device (100) separates metadata and unit object attribute information for the corresponding unit object using a delimiter to create a single string, repeats the string creation process for all unit objects within the polygon, and then creates documented data for the corresponding polygon.
[0081] The first language model can be trained using training data based on documented data about polygons. Therefore, the first language model can provide response data to user queries such as "What is the tallest building in Seoul?", "What is the largest apartment complex in Busan?", and "Please tell me the land cadastral information for the most expensive land in South Korea."
[0082]
[0083] FIG. 5 is a diagram illustrating a learning data generation process for spatial data based on a polyline according to one embodiment of the present disclosure.
[0084] As illustrated in FIG. 5, the computing device (100) inputs metadata including at least one piece of information from among coordinate information including the latitude and longitude coordinates of the center point of the polyline and the coordinates of the start and end points, structured address information obtained by geocoding the coordinate information, the line length of the polyline, the slope of the line, the average altitude of the line, or location characteristic information including surrounding environment information (distance to the sea, river, etc.) located in a preset change area based on the coordinate information, in the case of spatial data based on a polyline.
[0085] The computing device (100) separates metadata and unit object attribute information for the corresponding unit object with a delimiter to create a single string, repeats the string creation process for all unit objects within the polyline, and then creates documented data for the corresponding polyline.
[0086] The first language model can be trained using training data based on documented data about polylines. Therefore, the first language model can provide response data to user queries such as "What is the longest road name in Seoul?" and "Tell me a route that offers a view of the ocean."
[0087]
[0088] FIG. 6 is a diagram illustrating a learning data generation process for terrain data according to one embodiment of the present disclosure.
[0089] Referring to FIG. 6, when the location-based object data is terrain data, the computing device (100) sets a preset reference resolution as a unit column (S221). For example, when the computing device (100) sets the reference resolution to 10 cm and configures a unit column based on 1 m, 10 data per unit area (width × height) can be converted into text as one unit column.
[0090] The computing device (100) inputs one or more pieces of information as metadata among the center coordinate information of a coordinate group defined based on a set unit column, structured address information geocoded based on the center coordinate information, and location characteristic information including altitude, slope, slope direction, and land cover (mountain, river, road, etc.) for a change area set based on the center coordinate information (S222).
[0091] At this time, if the topographic data is contour data, the computing device (100) converts it into DEM data using a known DEM conversion tool and then inputs metadata based on the converted DEM data.
[0092] The computing device (100) separates metadata and unit object attribute information for the corresponding unit object using a delimiter to create a single string, repeats the string creation process for all unit objects in the terrain data, and then creates documented data for the corresponding terrain data.
[0093] The first language model can be trained using training data based on documented terrain data. Therefore, the first language model can provide response data to user queries such as "Find coniferous terrain with a slope of 30 degrees or more with a high risk of landslides."
[0094]
[0095] FIG. 7 is a diagram illustrating a learning data generation process for sensor data according to one embodiment of the present disclosure.
[0096] Referring to FIG. 7, in the case of sensor data, the computing device (100) communicates with IoT equipment and standard sensors to collect location information and measurement values of the sensors and inputs them as metadata (S231), and converts the metadata into text to generate documented data at preset time intervals (S232, S233).
[0097] At this time, the sensor may be an IoT sensor including at least one of a fine dust sensor for detecting the status of an object, a bridge vibration monitoring sensor, or a traffic volume measurement monitoring sensor.
[0098] The computing device (100) embeds documented data into a document file (document A, document B, document C, etc.) (S234) and stores the document file in a database (S235).
[0099] For example, the first language model can generate response data in real time using document files stored in a database for user queries such as "Where is the area with the worst fine dust right now?" or "Where is the area with the most traffic congestion in Seoul right now?"
[0100]
[0101] FIG. 8 is a diagram illustrating a learning data generation process for public data according to one embodiment of the present disclosure.
[0102] Referring to FIG. 8, when the location-based object data is public data such as weather information (typhoon, rain, weather, etc.) or performance information, the computing device (100) generates basic learning data based on a public API list including at least one public API (S241), and learns a first language model based on the basic learning data (S242).
[0103] When a user query is input (S243), the computing device (100) determines a public API matching the user query using a pre-learned first language model (S244), and calls the determined public API to receive a result value (S245, S246).
[0104] The computing device (100) converts the received result value into documented data and provides it as a prompt for the first language model (S246). Accordingly, the first language model returns response data to the user based on the prompt input (S247).
[0105] For example, the first language model can answer user queries such as "When do you think we will be affected by this typhoon?" or "What performance is on at the Seoul Arts Center tomorrow?" by searching for information matching the user query in real time.
[0106]
[0107] Meanwhile, if the location-based object data is spatial analysis data based on three-dimensional spatial information, the computing device (100) inputs spatial range information including the floor and lake for the object, and visibility information about the ratio of one or more environmental elements (sky, river, sea, forest, park, road, etc.) within the visible area of the object based on the result of the visibility analysis of the object as metadata.
[0108]
[0109] In this way, the present invention can convert location-based object data such as map information, sensor measurements, weather information, and performance information into documented data in text format and provide it as learning data for a generative artificial intelligence model.
[0110] Therefore, generative AI models can be retrained based on documented data on location-based object data, providing accurate answers to user queries that conventional generative AI cannot. For example, generative AI models can answer user queries related to location on a map based on accurate address information and coordinate information, such as building area, land area, and the mutual positional relationship between objects. Furthermore, they can provide accurate answers to queries about real-time data, such as weather, disaster relief, and sensor monitoring, based on real-time collected data.
[0111]
[0112] The various embodiments of the present invention described above can be combined with additional embodiments and modified within the scope understood by those skilled in the art in light of the detailed description above. It should be understood that the embodiments of the present invention are illustrative in all respects and not restrictive. For example, each component described as a single component may be implemented in a distributed manner, and likewise, components described as distributed may be implemented in a combined manner. Accordingly, all changes or modifications derived from the meaning, scope, and equivalent concepts of the claims of the present invention should be construed as being included within the scope of the present invention.
Claims
1. A method for generating learning data for generative artificial intelligence, performed by a computing device including at least one processor, A step of obtaining location-based object data for an object including a preset object; A step of converting metadata and object attribute data that specify the object based on the above location-based object data into a text listing format to create documented data; and A step of generating learning data based on the above documented data; Including, method.
2. In paragraph 1, The step of converting metadata and object attribute data that specify the object based on the above location-based object data into a text listing format and creating documented data is as follows. If the above object includes at least one unit object, a string is created by converting metadata and unit object attribute data that specify the unit object for each unit object into a text listing format, and the strings for each unit object are merged to create documented data. method.
3. In paragraph 1, The above location-based object data is, At least one of spatial data based on any one of points, polygons or polylines, topographic data based on digital elevation model (DEM) data, sensor data generated through real-time environmental monitoring, public data using an open application programming interface (API), or spatial analysis data based on three-dimensional spatial information. method.
4. In paragraph 3, The above metadata is, In the case where the above location-based object data includes spatial data based on the above point, it includes at least one piece of information from among coordinate information of the unit object, address information structured by geocoding the coordinate information, absolute altitude information of the coordinate information, and location characteristic information including average slope, slope direction, or distance to surrounding roads or major facilities for a change area set based on the coordinate information. method.
5. In paragraph 4, The step of converting metadata and object attribute data that specify the object based on the above location-based object data into a text listing format and creating documented data is as follows. The metadata converted into a text listing format and the unit object attribute data are separated by a delimiter to create a document string for each unit object. method.
6. In paragraph 3, The above metadata is, In the case where the above location-based object data includes spatial data based on the polygon, it includes at least one piece of information among structured address information by geocoding based on the central coordinate information of the polygon, and location characteristic information including the area, slope, slope direction or elevation above sea level of the polygon. method.
7. In paragraph 3, The above metadata is, In the case where the object data includes spatial data based on the polyline, it includes at least one of coordinate information including the latitude and longitude coordinates of the center point of the polyline and the coordinates of the start and end points, address information structured by geocoding the coordinate information, and location characteristic information including the line length of the polyline, the line slope or the average altitude of the line, or the surrounding environment information located in a preset change area based on the coordinate information. method.
8. In paragraph 3, The above metadata is, In the case where the object data includes the terrain data, the preset standard resolution is set as a unit column, and at least one of the following information is included: center coordinate information of a coordinate group defined based on the unit column, address information structured by geocoding based on the center coordinate information, and location characteristic information including elevation, slope, slope direction or land cover for a preset change area based on the center coordinate information. method.
9. In paragraph 3, The above metadata is, If the above location-based object data includes the sensor data, it includes the location information and measurement values of the sensor collected in preset time units. method.
10. In paragraph 9, The above sensor, An IoT sensor including at least one of a fine dust sensor, a bridge vibration monitoring sensor, or a traffic volume measurement monitoring sensor, method.
11. In paragraph 3, The step of converting metadata and object attribute data that specify the object based on the above location-based object data into a text listing format and creating documented data is as follows. If the above location-based object data includes public data, a step of generating basic learning data based on at least one public API; A step of learning a first language model based on the above basic learning data; A step of determining a public API matching a user query using a pre-trained first language model, and then calling the determined public API to receive a result value; and Including the step of converting the above result value into documented data and providing it as a prompt of the first language model. method.
12. In paragraph 3, The above metadata is, If the object data includes the spatial analysis data, at least one of the spatial range information including the layer and lake for the object or the visibility information for the ratio of one or more environmental elements based on the visibility analysis result of the object. method.
13. A computing device that generates learning data for generative artificial intelligence. a processor comprising at least one core; and A memory including program codes executable by the processor; The above processor, upon execution of the above program code, Obtain location-based object data for objects containing pre-defined objects, Based on the above location-based object data, metadata and object attribute data that specify the object are converted into text listing format to create documented data. Generating training data based on the above documented data, device.
Citation Information
Patent Citations
Battery protection device and battery pack including th same
KR1020240160419A
Method for providing digital drawing and digital drawing providing device
KR102129407B1
Method and System for Providing OTP to Integrated Emergency Broadcasting System
KR102479036B1
System and method for processing training data
KR102588531B1
Method and system for automatically annotating sensor data
WO2023135244A1