Data addressing method and device, equipment and storage medium
By generating and vectorizing data maps, constructing addressing vectors and knowledge maps, and using large language models to realize semantic data transformation, the problems of inaccurate data addressing and inability to achieve semantic data addressing in the prior art are solved, and efficient and accurate data addressing effect is achieved.
Patent Information
- Application Number
- CN202411999300.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately and quickly perform data addressing, cannot realize semantic data addressing, and the index data is one-sided, making it impossible to fully describe all data on the data page.
By generating a data map containing the entity data and relational data of the data page and vectorizing it, it constructs addressing vectors and knowledge graphs, and using large language models to achieve semantic data transformation and precise addressing.
Accurate semantic data addressing based on input information is realized, the system's addressing accuracy and efficiency are improved, and the problem of index data being one-sided and semantic data addressing cannot be achieved in traditional methods.
Smart Images

Figure CN119988641A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data retrieval technology, and in particular to a data addressing method, apparatus, device and storage medium. Background Art
[0002] Data addressing is a key technology in computer programming, determining how to efficiently access and manipulate data. In different application scenarios, the way data is addressed has a direct impact on performance.
[0003] Traditional technical approaches primarily rely on data ledgers and index matching. Data ledgers organize system addresses and content descriptions into a manual format, making it easier to query system addresses. Index matching uses database query statements to retrieve system addresses from a two-dimensional data table consisting of system addresses and indexes. However, both approaches have significant drawbacks. One-sided index data cannot accurately address the system. Due to storage limitations, index data must be concise and cannot fully describe all data on the data page, resulting in inaccurate system addressing. Both of these addressing technology paths rely on keyword retrieval, making semantic data addressing impossible. Summary of the Invention
[0004] The embodiments of the present application provide a data addressing method, apparatus, device, and storage medium to at least solve the technical problem in the related art that it is difficult to accurately and quickly address data.
[0005] According to one aspect of an embodiment of the present application, a data addressing method is provided, including:
[0006] Generate a data graph including data page entity data and data page relationship data based on the data page data, and perform vectorization processing on the data graph to generate a vectorized graph;
[0007] Vectorize the query information input by the user and construct an addressing vector;
[0008] Generate an addressing knowledge graph based on the addressing vector and the vectorized graph;
[0009] Utilizing the query information and the addressing knowledge graph, the data address matching the query information is output through the created addressing prompt words and large language model.
[0010] According to another aspect of an embodiment of the present application, a data addressing device is provided, comprising:
[0011] A knowledge graph construction module is used to generate a data graph containing data page entity data and data page relationship data based on data page data, and vectorize the data graph to generate a vectorized graph;
[0012] The query module is used to vectorize the query information input by the user and construct an addressing vector;
[0013] An addressing module, configured to generate an addressing knowledge graph based on the addressing vector and the vectorized graph;
[0014] An output module is used to use the query information and the addressing knowledge graph to output the data address that matches the query information through the created addressing prompt words and large language model.
[0015] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the data addressing method through the computer program.
[0016] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned data addressing method when running.
[0017] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0018] The data addressing method provided in the embodiment of the present application integrates data page data with system addresses and generates a data page knowledge graph based on comprehensive data page data. The knowledge graph is used as a data storage medium. During the query, the query information input by the user is also vectorized to construct an addressing vector and generate an addressing knowledge graph. By writing a prompt word template and referencing the large model's ability to understand text, the data semantic conversion is completed, thereby achieving precise system addressing based on the input information. By integrating the large language model and information retrieval technology, the addressing accuracy and efficiency of the system are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 is a flow chart of an optional data addressing method according to an embodiment of the present application;
[0021] Figure 2 is a schematic diagram of a data addressing method according to an embodiment of the present application;
[0022] Figure 3 is a schematic diagram of a data addressing method according to an embodiment of the present application;
[0023] Figure 4is a schematic diagram of a data addressing device according to an embodiment of the present application;
[0024] Figure 5 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] Currently, the existing data ledger addressing and index matching addressing technologies are both unable to accurately complete system addressing based on input information. The reasons are as follows:
[0028] 1. One-sided index data cannot accurately implement system addressing. Due to storage limitations, index data requires to be concise and cannot fully describe all the data on the data page, resulting in the inability to accurately implement system addressing.
[0029] 2. Dynamic data addressing cannot be achieved. Data pages will generate specific data page data content in different data time ranges and application examples. Dynamic data addressing cannot be achieved regardless of data ledger or index matching addressing.
[0030] 3. Multimodal data addressing cannot be achieved. Data ledger or index matching addressing can only realize text content retrieval and cannot address image data such as pictures and videos.
[0031] 4. Semantic data addressing cannot be achieved. Both of the above addressing technology paths use keyword retrieval, but cannot achieve semantic data addressing.
[0032] Based on this, this application proposes a data addressing method. Figure 1-3 The data addressing method of the embodiment of the present application is described in detail. Figure 1 As shown, the method mainly includes the following steps:
[0033] S101 generates a data graph including data page entity data and data page relationship data based on the data page data, performs vectorization processing on the data graph, and generates a vectorized graph.
[0034] In one embodiment, first, based on the data page data, data page holographic data, data page image sub-region data, and data address similarity relationships are generated.
[0035] Specifically, based on the data page data, generating data page holographic data includes: identifying the data page data based on optical character recognition technology to obtain data page multimodal data, where the multimodal data includes a data address and a data page context.
[0036] In one embodiment, optical character recognition (OCR) technology recognizes data page multimodal data Dt_MmMl. First, the text detection model CTPN is used to detect the text data location. CTPN can identify text areas in the image and determine their coordinate positions. Then, the CRNN model is used for text recognition. CRNN combines CNN and RNN, and can effectively recognize continuous text in the image. Then, the CTC model is used to output the serial number of the recognized text, and the output content includes: data address recognition Address_path and data page context Dt_Cont. Among them, the recognized data address Address_path contains the coordinate position and path information of the text, and Dt_Cont is a list of text content and its corresponding coordinates. The above outputs are integrated to generate data page multimodal data Dt_MmMl containing address information and text content, providing a basic data set for further data addressing and analysis.
[0037] Furthermore, using the data dictionary and prompt word template, data page labels, summaries and application scenarios are extracted and generated from the multimodal data; based on the data page labels, summaries, application scenarios and corresponding data addresses, data page holographic data is generated.
[0038] In one embodiment, the process of constructing data page holographic data (DtPage_HoLo) involves extracting and generating richer information from multimodal data (Dt_MmMl), and the specific steps are as follows:
[0039] Data Label Generation: Utilizing the contextual information (Dt_Cont) in the data page's multimodal data (Dt_MmMl) and the predefined data label dictionary (lbArr), the label generation function (fcreateLb) is called to create the data page label (DtPage_Label). The data label dictionary (lbArr) contains the label (lb) and the prompt word (prom), which guides the label generation function to generate the appropriate label based on the content.
[0040] Data page summary construction: Also based on the multimodal data Dt_MmMl, using the data summary prompt word template (Sum_Arr), the summary generation function (fcreateSummary) is called to generate the data page summary (DtPage_Summary). The summary generation function extracts key information based on the data content and prompt word template to form a summary.
[0041] Data Page Application Scenario Construction: Again, using the Dt_MmMl data and the data page application scenario dictionary (sceArr), the scenario generation function (fcreateSce) is called to create the data page application scenario (DtPage_Scenario). The application scenario dictionary (sceArr) contains a prompt word (prom), which guides the scenario generation function to generate an application scenario based on the content.
[0042] Data encapsulation: Combine the generated data label (DtPage_Label), summary (DtPage_Summary) and application scenario (DtPage_Scenario) with the original data address (Address_path) and encapsulate them into a data list.
[0043] The final generated data page holographic data (DtPage_HoLo) contains the label, summary and application scenario corresponding to each data address, providing a comprehensive description of the data page.
[0044] This process converts the original multimodal data into data page holographic data containing rich metadata by defining and using dictionaries, templates and generation functions, which facilitates subsequent data management and analysis.
[0045] Furthermore, based on the data page data, data page image sub-area data is generated, including: spatially dividing the data page into sub-areas by coordinate annotation to generate image sub-areas; applying the coordinates of the image sub-areas to perform spatial calculations on the multimodal data to obtain the data contained in the image sub-areas, and applying the prompt word function to generate a sub-area data summary; and obtaining the data page image sub-area data based on the data contained in the image sub-area and the sub-area data summary.
[0046] In one embodiment, the process of constructing the data page image sub-region data (DtImg_Data) involves dividing the data page into multiple sub-regions and generating summary information for each sub-region.
[0047] First, image sub-region division is performed: the data page is divided into multiple image sub-regions by coordinate annotation. Each sub-region is defined by four coordinate values: (x1, y1, x2, y2), which represent the positions of the upper left corner and lower right corner of the sub-region.
[0048] Sub-area data extraction: Using these coordinates, perform spatial calculations on the generated multimodal data (Dt_MmMl) to extract the data content contained in each sub-area.
[0049] Sub-area content aggregation: Call the sub-area content aggregation function (fareasum) to aggregate the extracted data content (Dt_Cont) with the corresponding sub-area coordinates (area^iy) to generate sub-area aggregation data (areaCont^iy).
[0050] Sub-area data encapsulation: Encapsulate the aggregated data of each sub-area (areaCont^iy) into a data list (ImgDt_Area), each list item contains the address path (address_path) and the aggregated content of all sub-areas under the path.
[0051] Sub-region summary generation: Define the summary prompt word (summaryProm) and call the summary generation function (fSummary) to generate a summary (arrSummary^iy) for the data content of each sub-region. The generated sub-region summary (arrSummary^iy) is encapsulated with the corresponding address path (Address_path) to form the final image sub-region data (DtImg_Data).
[0052] This process generates detailed data summaries for each image sub-area of the data page through spatial calculation and content aggregation. These summaries can be used for subsequent data analysis and retrieval tasks, improving the efficiency of data management and utilization.
[0053] Furthermore, based on the data page data, a data address similarity relationship is generated, including: based on the data address and the sub-area data summary, calling a similarity calculation function to calculate the similarity of each image sub-area, screening out associated addresses with a similarity higher than a threshold based on the similarity of each image sub-area, and obtaining the image similarity of the associated addresses.
[0054] In one embodiment, a similarity algorithm and a calculation function fSimilarity are used in conjunction with the generated image subregion data (DtImg_Data) to calculate the similarity between the images on each data page, generating a temporary similarity result temp_Similarity^i. The similarity-related addresses with similarity values above the average are filtered using the avgFilter function to extract the similarity-related addresses, forming ImgDt_Similarity^i.
[0055] Based on the data page label, summary and application scenario, the similarity calculation function is called to calculate the similarity between the holographic data of each data page. Based on the similarity of the holographic data of each data page, the associated addresses with similarity higher than the threshold are filtered out, and the holographic data similarity of the associated addresses is obtained.
[0056] In one embodiment, the generated data page holographic data (DtPage_HoLo), including data tags, summaries, and application scenarios, is used to call the fSimilarity function again to calculate the similarity between the holographic data to generate temp_Similarity^i. Similarly, the avgFilter function is used to filter and extract similarity-related addresses with similarity values higher than the average to form DtPage_Similarity^i.
[0057] Based on the image similarity of the associated address, the holographic data similarity and the preset weight, a weighted sum is performed to obtain the data address similarity relationship.
[0058] Specifically, we define the weights for image similarity (addr_wg) and holographic data similarity (dtPage_wg). Multiply the similarity values of ImgDt_Similarity^i and DtPage_Similarity^i by their corresponding weights, then sum them to calculate the final data address similarity relationship, Addr_RRship^i. This relationship combines image and content similarity information to represent the similarity between different data addresses.
[0059] This process provides a comprehensive method to evaluate the similarity relationship between different data addresses by combining the similarity calculation of image sub-area data and holographic data, as well as weight adjustment.
[0060] Furthermore, entity data and relationship data are created based on the data page holographic data, data page image sub-area data, and data address similarity relationships. Entity data includes address entities, label entities, summary entities, and scene entities. Relationship data includes address-to-address relationships, address-to-label relationships, address-to-summary relationships, and address-to-scene relationships. A data graph is generated based on the entity data and relationship data.
[0061] Furthermore, entity data and relationship data are created based on the data page holographic data, data page image sub-area data, and data address similarity relationships. The entity data includes a data page address entity, a data page label entity, a data page summary entity, and a data page application scenario entity. The relationship data includes the relationship between data page addresses and data page addresses, the relationship between data page addresses and data page labels, the relationship between data page addresses and data page summaries, and the relationship between data page addresses and data page application scenarios. A data graph is generated based on the data page entity data and data page relationship data.
[0062] In one embodiment, the fentity function is used to create different types of graph entities based on the provided data page holographic data, data page image sub-area data, and data address similarity relationships. Entities include addresses (Kg_Eng[Addr]), labels (Kg_Eng[Label]), summaries (Kg_Eng[Summary]), and scenarios (Kg_Eng[Scenario]).
[0063] Use the relationship function to create relationships between entities based on the main entity (mainEntity), related entity (entity), and type (type). Relationships include those between addresses and labels (Kg_AddrR[Label]), summaries (Kg_AddrR[Summary]), scenarios (Kg_AddrR[Scenario]), and similar addresses (Kg_AddrR[Sim]).
[0064] Each relationship contains a unique identifier (id), a path identifier (pathid), related entity identifiers (such as labelid, summaryid, etc.), and a relationship name (name).
[0065] The resulting data page data graph (Page_kg) consists of data page entity data (Kg_Eng[]) and data page relationship data (Kg_AddrR[]), forming a structured graph that can be used to represent the structure and content of the data page, as well as the relationships between entities. This makes the relationships between data clearer and facilitates data management and retrieval. The graph allows for efficient data organization and query, improving the efficiency and accuracy of data processing.
[0066] Furthermore, the data graph is vectorized to generate a vectorized graph.
[0067] In one embodiment, the fcollection function is used to construct a vector dataset based on graph entities (such as labels, summaries, and scenes) and relationships (such as similarity), as well as their corresponding names (such as label, summary, scene, sim, and cont). This function returns a tuple containing vector representations of different types of entities, such as addresses (Coll_[Addr]Embedding), labels (Coll_[Label]Embedding), summaries (Coll_[Summary]Embedding), and scenes (Coll_[Scen]Embedding).
[0068] The Two-Point Technique algorithm is applied to deduplicate the vector dataset to ensure the uniqueness of each entity and relationship in the vector space. This step compares the vectors in the vector dataset and removes duplicate or similar vectors to avoid introducing redundancy into the graph.
[0069] The deduplicated vector datasets are merged to form the final vectorized graph (Kg_Vector). This vectorized graph contains vector representations of all entities and relationships in the data graph, providing the basis for subsequent similarity calculations, graph queries, and data analysis.
[0070] Through this process, graph data can be converted into a vector form suitable for machine learning and data analysis, making it possible to use vector space models to explore complex patterns and connections between entities and relationships. This approach improves the operability and analysis efficiency of graph data, especially when processing large-scale graph data.
[0071] S102 performs vectorization processing on the query information input by the user to construct an addressing vector.
[0072] In one embodiment, query information input by a user is vectorized to construct an addressing vector, including: vectorizing the query information to generate vectorized query information; based on the vectorized query information, using a vector distance algorithm to query in a vectorized graph to obtain an initial vector result; and deduplicating paths in the initial vector result to obtain an addressing vector.
[0073] Specifically, the input information (Input_Info) is converted into a vector form through vectorization (Vector Quantization) processing to generate vectorized input information (Input_Info_Vector).
[0074] Using vector distance algorithms (such as cosine similarity and Euclidean distance), we construct a vector query algorithm (queryVector) to search the generated vectorized graph (Kg_Vector) for the vector that best matches the input vector. We use the Two-Pointer Technique algorithm to deduplicate paths in the query result (temp_vector) to ensure that each path appears only once.
[0075] The final deduplicated vector result (query_vector) contains the collection name (collname), identifier (id), and value (value), which are used for subsequent addressing operations.
[0076] This process generates an accurate addressing vector by converting the input information into vector form and performing queries and deduplication in the vectorized graph, providing the system with an efficient way to accurately address the input query information.
[0077] S103 generates an addressing knowledge graph based on the addressing vector and the vectorized graph.
[0078] In one embodiment, an addressing knowledge graph is generated based on an addressing vector and a vectorized graph, including: extracting a graph ID set from the addressing vector, and using a query language to construct an initial addressing knowledge graph based on the graph ID set; obtaining the associated nodes of each node in the initial addressing knowledge graph based on the similarity relationship, merging the initial addressing knowledge graph with the associated nodes, and generating an addressing knowledge graph.
[0079] Specifically, a set of graph IDs (kg_ids) is extracted from the generated query vector set (query_vector), and these IDs correspond to related entities in the vectorized graph.
[0080] Using the Cypher query language, according to the graph ID set (kg_ids), a matching query (match) is performed to build a preliminary addressing knowledge graph (kg'), including entities (a, b) and the relationship (r) between them.
[0081] Based on the similarity relationship (Addr_RRship), Cypher query language is used to obtain the associations (kg”) with a depth of 1 directly connected to each entity. These associations reflect the similarity relationship between entities.
[0082] The preliminary addressing knowledge graph (kg') is merged with the association graph (kg") of depth 1 to generate the final addressing knowledge graph (AddrKG).
[0083] This process combines addressing vectors with a vectorized graph to construct an addressable knowledge graph containing entities and relationships. This graph provides detailed data addresses and contextual information related to the input information. This graph can be used to support complex query and data analysis tasks, especially in scenarios that require understanding complex relationships between data.
[0084] S104 utilizes the query information and the addressing knowledge graph, and outputs the data address that matches the query information through the created addressing prompt words and large language model.
[0085] In one embodiment, query information and an addressing knowledge graph are utilized to output data addresses that match the query information through created addressing prompt words and a large language model, including: creating addressing prompt words, the addressing prompt words including the addressing knowledge graph, the query information, and a prompt message, the prompt message being used to indicate the generation of matching data addresses based on the query information and the addressing knowledge graph, and providing reasons and ranking; inputting the addressing prompt words and the large language model into a preset address generation function, and outputting a list of addresses that match the query information and the reasons for the matching.
[0086] Specifically, a prom is defined based on the generated addressing knowledge graph (AddrKG) and input information (Input_Info). This prom contains the addressing knowledge graph data, query information, and a specific prompt message (promMsg). The prompt message indicates that the system needs to generate the most matching data address based on the input information (Input_Info) and the addressing knowledge graph (AddrKG) data, and provides the reasoning and ranking of the match.
[0087] Furthermore, using the defined addressing hint word (prom) and the Large Language Model (LLM), the address generation function (ftopaddr) is called to process the addressing request. The function (ftopaddr) processes the information in the hint word and searches the addressing knowledge graph for the data addresses most relevant to the input information, generating a ranked list of addresses (TopAddr). The address list (TopAddr) is returned in JSON format, containing the most matching data addresses and the corresponding matching reasons, making it easier for users to understand and use.
[0088] This process achieves precise data addressing based on input information by intelligently utilizing knowledge graphs and natural language processing technologies, thereby improving the efficiency and accuracy of data retrieval.
[0089] The data addressing method of the embodiment of the present application is based on a large language multimodal model and optical character recognition (OCR) as basic technologies, uses a knowledge graph as a data storage medium, and integrates a large language model and information retrieval technology to achieve system addressing according to input information.
[0090] First, optical character recognition (OCR) technology is used to identify multimodal data, including data address data and data page context. Then, large-scale model technology is used to construct data page holographic data. By defining image sub-area data, dynamic sub-area data is generated. Then, based on the data page holographic data and dynamic sub-area data, a data address association is established using a similarity algorithm. Based on the identified data and address relationships, a data page graph is constructed, including data entity data and relationship data. The graph data is then vectorized and stored to generate a vectorized graph. Finally, based on the input query, prompt words are written, and the large-scale model's ability to understand text is utilized to complete the semantic conversion of the data page holographic data. This allows the system to accurately address the input information and return an address list.
[0091] In order to facilitate understanding of the data addressing method of the embodiment of the present application, the following Figure 2 and 3 Further description.
[0092] like Figure 2 As shown, the data addressing method includes identifying multimodal data, building a data graph, and input (query requirement) addressing.
[0093] like Figure 3 As shown, identifying multimodal data includes identifying data page multimodal system data Dt-MMML through image recognition models such as OCR technology, further constructing data page holographic data DtPage-HoLo, constructing image sub-area data Dtlmg-Data, and constructing the association relationship Addr_Rrship of each data address.
[0094] Furthermore, the data page data map Page_kg and the map vectorization Kg_Vector are constructed.
[0095] Furthermore, the addressing vector query_vector is constructed, the addressing knowledge graph AddrKG is constructed, and the addressing address TopAddr is generated.
[0096] This application achieves multiple technical breakthroughs and beneficial effects in system addressing by comprehensively applying OCR technology, large language model (LLM), multimodal data processing and knowledge graph:
[0097] Comprehensive Data Information Construction: This application uses optical character recognition (OCR) technology to dynamically identify system addresses and contextual information as foundational data. Using the Large Language Model (LLM), it constructs holographic data for data pages, generating comprehensive data information. This provides more comprehensive data information than traditional indexing, addresses the inaccurate addressing issues caused by one-sided index data, and enhances addressing accuracy and reliability.
[0098] Implementation of dynamic data retrieval: This application can obtain data corresponding to image sub-areas by spatially dividing data pages. Through the construction of image sub-area data and the acquisition of dynamic data, this application can realize real-time retrieval of active data, solving the problem that traditional methods cannot effectively retrieve dynamic data.
[0099] Multimodal data addressing: By adopting multimodal data recognition technology, integrating image information data, data page holographic data and system address, and using knowledge graph as storage medium, this application realizes multimodal data addressing and expands the application scope of addressing technology.
[0100] Precise semantic addressing: When searching, this application can complete the semantic conversion of data by writing addressing prompt words and utilizing the text understanding ability of the large language model (LLM), thereby achieving precise addressing based on the input information and improving the semantic accuracy of the addressing.
[0101] In summary, the beneficial effects of this application include improving the accuracy, dynamism and multimodality of addressing, while also enhancing the system's ability to process semantic information, providing users with a more powerful and flexible system addressing solution.
[0102] According to another aspect of the embodiment of the present application, a data addressing device for implementing the above data addressing method is also provided. Figure 4 As shown, the device includes:
[0103] The knowledge graph construction module 401 is used to generate a data graph including data page entity data and data page relationship data based on the data page data, and perform vectorization processing on the data graph to generate a vectorized graph;
[0104] Query module 402, used to vectorize the query information input by the user and construct an addressing vector;
[0105] An addressing module 403 is configured to generate an addressing knowledge graph based on the addressing vector and the vectorized graph;
[0106] The output module 404 is used to use the query information and the addressing knowledge graph to output the data address that matches the query information through the created addressing prompt words and large language model.
[0107] It should be noted that the data addressing device provided in the above embodiment, when executing the data addressing method, is merely illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data addressing device provided in the above embodiment and the data addressing method embodiment are based on the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0108] According to another aspect of the embodiments of the present application, an electronic device corresponding to the data addressing method provided in the above embodiments is also provided to execute the above data addressing method.
[0109] Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. Figure 5 As shown, the electronic device includes: a processor 500, a memory 501, a bus 502 and a communication interface 503, and the processor 500, the communication interface 503 and the memory 501 are connected via the bus 502; the memory 501 stores a computer program that can be run on the processor 500, and when the processor 500 runs the computer program, it executes the data addressing method provided in any of the aforementioned embodiments of the present application.
[0110] The memory 501 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The system network element and at least one other network element are connected via at least one communication interface 503 (which may be wired or wireless), and may use the Internet, a wide area network, a local area network, a metropolitan area network, or the like.
[0111] Bus 502 may be an ISA bus, a PCI bus, or an EISA bus. Buses may be classified as address buses, data buses, and control buses. Memory 501 is used to store programs, and processor 500 executes the programs upon receiving execution instructions. The data addressing method disclosed in any of the aforementioned embodiments of the present application may be applied to processor 500 or implemented by processor 500.
[0112] The processor 500 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 500 or by software instructions. The above processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 501 , and the processor 500 reads the information in the memory 501 and completes the steps of the above method in combination with its hardware.
[0113] The electronic device provided in the embodiment of the present application and the data addressing method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.
[0114] According to another aspect of the embodiments of the present application, a computer-readable storage medium corresponding to the data addressing method provided in the aforementioned embodiments is also provided, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the data addressing method provided in any of the aforementioned embodiments.
[0115] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0116] The computer-readable storage medium provided by the above-mentioned embodiments of the present application and the data addressing method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0117] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0118] The above embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A data addressing method, characterized in that: include: Generate a data graph including data page entity data and data page relationship data based on the data page data, and vectorize the data graph to generate a vectorized graph; Vectorize the query information input by the user and construct an addressing vector; Based on the addressing vector and the vectorized graph, generate an addressing knowledge graph; Utilizing the query information and the addressing knowledge graph, through the created addressing hint words and large language model, the data address matching the query information is output.
2. The method according to claim 1, characterized in that Generate a data graph containing data page entity data and data page relationship data based on data page data, including: Based on the data page data, generating data page holographic data, data page image sub-area data and data address similarity relationship; Creating entity data and relationship data according to the data page holographic data, the data page image sub-area data and the data address similarity relationship, wherein the entity data includes an address entity, a label entity, a summary entity and a scene entity, and the relationship data includes an address-address relationship, an address-label relationship, an address-summary relationship and an address-scene relationship; The data graph is generated based on the entity data and the relationship data.
3. The method according to claim 2, characterized in that Based on the data page data, generating the data page holographic data includes: Recognize the data page data based on optical character recognition technology to obtain data page multimodal data, wherein the multimodal data includes a data address and a data page context; Extracting and generating data page labels, summaries, and application scenarios from the multimodal data using a data dictionary and a prompt word template; The data page holographic data is generated based on the data page tag, summary, application scenario and data address.
4. The method according to claim 3, characterized in that Generating the data page image sub-area data based on the data page data includes: Dividing the data page into spatial sub-areas in the form of coordinate marking to generate image sub-areas; Using the coordinates of the image sub-region to perform spatial calculation on the multimodal data, obtaining the data contained in the image sub-region, and using a prompt word function to generate a sub-region data summary; The data page image sub-area data is obtained according to the data contained in the image sub-area and the sub-area data summary.
5. The method according to claim 4, characterized in that Generating the data address similarity relationship based on the data page data includes: Based on the data address and the sub-area data summary, calling a similarity calculation function to calculate the similarity of each image sub-area, filtering out associated addresses with similarities higher than a threshold based on the similarities of each image sub-area, and obtaining image similarities of the associated addresses; Based on the data page label, summary and application scenario, calling a similarity calculation function, calculating the similarity between the holographic data of each data page, filtering out the associated addresses with a similarity higher than a threshold based on the similarity of the holographic data of each data page, and obtaining the holographic data similarity of the associated addresses; Based on the image similarity, holographic data similarity and preset weights of the associated addresses, a weighted sum is performed to obtain the data address similarity relationship.
6. The method according to claim 1, characterized in that Vectorize the query information input by the user and construct an addressing vector, including: Vectorizing the query information to generate vectorized query information; Based on the vectorized query information, use a vector distance algorithm to query the vectorized graph to obtain an initial vector result; De-duplication is performed on the paths in the initial vector result to obtain the addressing vector.
7. The method according to claim 1, characterized in that Based on the addressing vector and the vectorized graph, an addressing knowledge graph is generated, including: Extracting a graph ID set from the addressing vector, and constructing an initial addressing knowledge graph according to the graph ID set using a query language; According to the similarity relationship, the associated nodes of each node in the initial addressing knowledge graph are obtained, and the initial addressing knowledge graph and the associated nodes are merged to generate the addressing knowledge graph.
8. The method according to claim 1, characterized in that Using the query information and the addressing knowledge graph, the data address matching the query information is output through the created addressing prompt words and large language model, including: Creating an addressing prompt word, wherein the addressing prompt word includes an addressing knowledge graph, query information, and a prompt message, wherein the prompt message is used to indicate that a matching data address is generated according to the query information and the addressing knowledge graph, and to provide reasons and rankings; The addressing prompt words and the large language model are input into a preset address generation function, and a list of addresses matching the query information and matching reasons are output.
9. A data addressing device, characterized in that: include: A knowledge graph construction module, used to generate a data graph including data page entity data and data page relationship data based on data page data, and vectorize the data graph to generate a vectorized graph; The query module is used to vectorize the query information input by the user and construct an addressing vector; An addressing module, used for generating an addressing knowledge graph based on the addressing vector and the vectorized graph; The output module is used to utilize the query information and the addressing knowledge graph to output the data address matching the query information through the created addressing prompt words and large language model.
10. An electronic device, characterized in that: The invention comprises a processor and a memory storing program instructions, wherein the processor is configured to execute the data addressing method according to any one of claims 1 to 8 when executing the program instructions.
11. A computer-readable medium, characterized in that Computer-readable instructions are stored thereon, and the computer-readable instructions are executed by a processor to implement a data addressing method as claimed in any one of claims 1 to 8.
Citation Information
Cited By
Addressing matrix construction method based on multi-dimensional identification
CN120750906A