Method and device for analyzing unstructured data and storage medium

The automated execution plan is built through large language models and semantic operator matching, which solves the problems of complex query automation and dynamic adjustment in unstructured data analysis, improves analysis efficiency and accuracy, and is suitable for unstructured data analysis of non-professional users.

CN120578740AActive Publication Date: 2025-09-02TSINGHUA UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510726851.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-27
Filing Date
2025-05-30
Publication Date
2025-09-02
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The prior art has problems in unstructured data analysis that are difficult to automatically process complex queries, high cost to build complex analysis frameworks, and lack of dynamic adjustment capabilities, resulting in the analysis results being easily distorted.

Method used

Semantic analysis is performed through large language models, logical representation is generated, candidate operator matching and reordering are combined with predefined semantic operators, automated execution plans are built, and parallel execution strategies of topological sorting are adopted to dynamically adjust the execution process to achieve automation and real-time optimization of complex queries.

Benefits of technology

It enables non-professional users to easily perform complex unstructured data analysis, improves analysis efficiency and result accuracy, solves the shortcomings of traditional methods in multi-step reasoning and dynamic adjustment, and ensures the reliability of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578740A_ABST
    Figure CN120578740A_ABST
Patent Text Reader

Abstract

The invention provides a method and equipment for analyzing unstructured data and a storage medium. The method comprises the following steps: constructing a predefined standard semantic operator set, wherein each operator has logic representation; receiving a natural language query input by a user as a to-be-simplified query; performing semantic analysis on the to-be-simplified query by using the large language model to generate a logic representation of the to-be-simplified query; candidate operators are matched and screened based on the semantic similarity between the logic representation and the operator logic representation; the candidate operators are reordered; sequentially applying candidate operators to simplify the query until the simplification is successful; and when it is determined that the current simplified query is a non-redecomposable minimum semantic unit, constructing and executing a plan according to a query simplification process. Through automatic plan generation and optimization, the technical threshold of unstructured data analysis is remarkably reduced, and the query efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing and analysis, and in particular to a method, device, and storage medium for analyzing unstructured data. Background Art

[0002] Data lake technology is widely used for its ability to process large amounts of heterogeneous data. However, intelligent analysis, particularly of unstructured data, remains challenging, as existing tools like SQL offer limited support. The primary goal of unstructured data analysis is to quickly and accurately extract valuable information. Efficiently supporting complex semantic analysis has become a research focus, but existing solutions present high technical barriers for average users and have limited effectiveness.

[0003] Specifically, although some systems allow users to create analysis processes through coding, they place excessive demands on users' technical capabilities. Methods based on retrieval-augmented generation (RAG) have a narrow scope of application and can only handle small amounts of data queries with simple logic, and their accuracy is not high. At the same time, existing large-scale language model applications rely heavily on static, manually orchestrated execution pipelines when processing complex queries. These pipelines are not only complex and costly to build, but also lack the ability to dynamically adapt to intermediate errors, making analysis results easily distorted. Traditional RAG methods are also insufficient in multi-step reasoning and aggregate analysis.

[0004] Therefore, the main challenge of current unstructured data analysis lies in how to automatically process complex queries and build an analysis framework that can automatically generate plans, dynamically execute and adjust in real time. Summary of the Invention

[0005] The purpose of this application is to provide a method for analyzing unstructured data, which can realize automatic plan generation and orchestration for complex natural language queries, and can also realize dynamic and interactive execution and dynamic adjustment based on real-time results.

[0006] According to one embodiment of the present application, a method for analyzing unstructured data is proposed, comprising:

[0007] receiving a natural language query input by a user as a query to be simplified;

[0008] Use a large language model to perform semantic analysis on the query to be simplified and generate a logical representation of the query to be simplified;

[0009] Matching the logical representation of the query to be simplified with the logical representations of a plurality of predefined standard semantic operators based on semantic similarity, screening out candidate operators, each of the standard semantic operators having at least one logical representation;

[0010] Evaluate the degree and executability of the candidate operators in solving the query to be simplified, and re-rank the candidate operators;

[0011] Apply the candidate operators to simplify the query in sequence according to the sorting results until the query is successfully simplified to obtain the simplified query.

[0012] When it is determined that the currently simplified query is a minimum semantic unit that cannot be decomposed any further, an execution plan is constructed according to the query simplification process, and the execution plan is executed using a parallel execution strategy based on topological sorting, and an analysis result is output.

[0013] In some embodiments, the plurality of standard semantic operators include some or all of the following operators: Retrieve, Filter, Aggregate, Validate, OrderBy, Compare, Refine, Translate, Transform, Scan, Extract, Conceptualize, Generate, Explain, GroupBy, Cluster, Classify, Link, SetOP, Integrate, Compute, and Summarize.

[0014] In some embodiments, the logical representation of the standard semantic operator is a structured template obtained by replacing the specific values ​​of the semantic elements in the natural language fragment with corresponding placeholders, which is used to represent the logical meaning of the natural language fragment. The placeholders include entity placeholders, condition placeholders, and attribute placeholders.

[0015] In some embodiments, each of the standard semantic operators has at least one corresponding physical operator, which is implemented by a pre-programmed algorithm or by processing semantic operations using a large language model. Some or all of the multiple standard semantic operators have both physical operators implemented by pre-programmed algorithms and physical operators implemented by processing semantic operations using a large language model.

[0016] In some implementations, a large language model is used to perform semantic parsing on the query to be simplified to generate a logical representation of the query to be simplified, including:

[0017] Providing prompt words to the large language model to instruct the large language model to parse semantic elements from the query to be simplified, wherein the semantic elements include entities, conditions, and attributes;

[0018] The large language model replaces the specific values ​​corresponding to entities, conditions, and attributes in the query to be simplified with corresponding entity placeholders, condition placeholders, and attribute placeholders to generate a structured logical representation.

[0019] In some embodiments, candidate operators are screened based on semantic similarity between the logical representation of the query to be simplified and the logical representations of a plurality of predefined standard semantic operators, including:

[0020] Convert the logical representation of the query to be simplified into a query semantic embedding vector;

[0021] Using the query semantic embedding vector as the query vector, a nearest neighbor search is performed in a pre-built vector index to calculate the similarity between the query semantic embedding vector and the operator semantic embedding vector, wherein the vector index stores the operator semantic embedding vector converted from the logical representation of the standard semantic operator;

[0022] The standard semantic operators corresponding to the semantic embedding vectors of multiple operators with the highest similarity are selected as candidate operators.

[0023] In some embodiments, evaluating the degree to which candidate operators solve the query to be simplified and their executability, and reordering the candidate operators, includes:

[0024] Use a large language model to check whether the preconditions of each candidate operator have been met;

[0025] Using a large language model, evaluate the degree to which each candidate operator that meets the preconditions solves the query to be simplified, which can be completely solved, partially solved, or not solved at all;

[0026] The candidate operators are ranked based on the degree to which they solve the query to be simplified and the semantic similarity between the logical representation of the candidate operators and the logical representation of the query to be simplified.

[0027] In some implementations, candidate operators are sequentially applied to simplify the query to be simplified according to the ranking results until the current query to be simplified is successfully simplified to obtain a simplified query, including:

[0028] Using a large language model, the expected outputs of the re-ranked candidate operators are sequentially applied to replace the matching fragments of the candidate operators in the query to be simplified, and the query after replacement is judged to see whether it meets the preset criteria.

[0029] The candidate operator of the query that meets the preset criteria obtained after the first replacement is regarded as the successfully applied operator, and the corresponding query that meets the preset criteria is regarded as the simplified query;

[0030] Generate intermediate result data and corresponding natural language descriptions representing the execution results of successfully applied operators, and add them to the set of available variables for use by subsequent operators.

[0031] In some implementations, the method further includes, after obtaining the simplified query, determining whether the simplified query is a minimum semantic unit that cannot be further decomposed by:

[0032] Using a large language model, it is determined whether the simplified query is a minimal semantic unit that cannot be decomposed any further. The minimal semantic unit includes a single entity, a numerical value, or a Boolean condition.

[0033] In some embodiments, the method further comprises:

[0034] When it is determined that the current simplified query is not the smallest semantic unit that cannot be decomposed any further, the current simplified query is used as a new query to be simplified, and matching based on semantic similarity, reordering of candidate operators, and simplification of the query to be simplified are re-executed until it is determined that the current simplified query is the smallest semantic unit that cannot be decomposed any further.

[0035] In some implementations, executing the execution plan using a parallel execution strategy based on topological sorting includes:

[0036] Topological sorting is performed based on the dependencies between operators, and operators whose preconditions are met and have no dependencies on each other are executed in parallel.

[0037] In some embodiments, the method further includes, during the execution of the execution plan, if the intermediate result data generated does not meet expectations, adjusting the execution plan based on the current execution plan, regenerating the execution plan, or rewriting the original natural language query to re-identify relevant data.

[0038] In some embodiments, the method further includes the steps of semantically parsing and preprocessing the unstructured data, including:

[0039] Perform word segmentation, entity recognition, and semantic annotation on the input unstructured data to generate an initial semantic feature vector;

[0040] Utilize a large language model to perform contextual analysis on the initial semantic feature vector to obtain an enhanced context-dependent semantic representation.

[0041] The enhanced context-sensitive semantic representation is stored in a structured machine-readable intermediate format.

[0042] According to one embodiment of the present application, an electronic device is proposed, which includes a memory and a processor, wherein the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement any of the methods described above when executing the computer instructions.

[0043] According to one embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in any one of the above items is implemented.

[0044] According to the solution for analyzing unstructured data provided by this application, the analysis requirements input by users through a natural language query interface can be accepted. After receiving the query, the query can first be semantically parsed with the assistance of a large language model to generate a structured logical representation. Subsequently, combined with predefined semantic analysis operators, the query execution plan is automatically constructed and optimized through steps such as automated operator matching, candidate operator reordering, and recursive query simplification. This structured processing method avoids dependence on the user's advanced programming skills, allowing people with non-professional backgrounds to easily perform complex unstructured data analysis tasks, and significantly improves the overall efficiency and accuracy of mining valuable information from large-scale unstructured data, effectively overcoming the problems of low efficiency and insufficient precision of traditional methods in this regard.

[0045] Moreover, the solution for analyzing unstructured data provided by this application can understand and process queries involving multi-step reasoning and complex logical analysis, surpassing the limitations of traditional RAG and other methods that can only process simple queries. According to this application, the query is recursively simplified into the smallest semantic unit to construct a logical plan, and is efficiently executed through a parallel execution strategy based on topological sorting. In addition, during the execution process, according to this application, the intermediate results can be verified, and when it is found that they do not meet expectations, self-correction and optimization can be performed by rewriting the query, re-identifying the data, and dynamically adjusting the execution plan, etc., which solves the problem that the traditional static analysis pipeline cannot cope with the failure or uncertainty of intermediate operations, thereby ensuring the reliability of the analysis results.

[0046] Other features and advantages of the technical solution proposed in this application are described in detail below. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.

[0048] Figure 1 A flow chart of a method for analyzing unstructured data according to an embodiment of the present application is shown.

[0049] Figure 2The diagram shows an architecture diagram of a plan generation system according to an exemplary embodiment of the present application.

[0050] Figure 3 An exemplary diagram of an automated plan generation method according to an exemplary embodiment of the present application is shown.

[0051] Figure 4 A schematic diagram of semantic parsing of a natural language query according to an exemplary embodiment of the present application is shown.

[0052] Figure 5 A schematic diagram of matching a suitable operator for a query based on semantic matching according to an exemplary embodiment of the present application is shown.

[0053] Figure 6 A schematic diagram of reordering candidate operators according to an exemplary embodiment of the present application is shown.

[0054] Figure 7 A schematic diagram of applying candidate operators to simplify queries according to an exemplary embodiment of the present application is shown.

[0055] Figure 8 A schematic diagram of determining whether a query is completely decomposed according to an exemplary embodiment of the present application is shown.

[0056] Figure 9 A schematic diagram of a complete query decomposition and plan generation according to an exemplary embodiment of the present application is shown.

[0057] Figure 10 A flowchart of plan execution according to an exemplary embodiment of the present application is shown.

[0058] Figure 11 It is a structural diagram of an electronic device shown in at least one embodiment of the present application. DETAILED DESCRIPTION

[0059] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0060] The embodiments of the present application can be applied to a computer system / server that can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the computer system / server include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above, among others.

[0061] Computer systems / servers may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, and the like, that perform specific tasks or implement specific abstract data types. Computer systems / servers may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located on local or remote computer system storage media, including storage devices.

[0062] This application provides a method for analyzing unstructured data. Figure 1 The method mainly includes the following steps 1 to 6.

[0063] Step 1: Receive a natural language query input by a user as a query to be simplified.

[0064] Users can use the system's query interface to input their analysis requirements for unstructured data, typically expressed in natural language. For example, a user might enter, "Among questions with over 500 views, which sport involving a ball has the highest ratio of injury-related questions to training-related questions?" The system receives this raw natural language query and uses it as the initial query to be simplified, preparing for subsequent parsing and processing.

[0065] Step 2: Use a large language model to perform semantic analysis on the query to be simplified and generate a logical representation of the query to be simplified.

[0066] Step 2 can utilize the capabilities of a large language model to convert the query to be simplified (eg, a query in natural language or containing natural language fragments) into a structured logical representation that is easier for a machine to understand and process.

[0067] In some implementations, the semantic parsing process may include:

[0068] A prompt is provided to the large language model to instruct it to parse semantic elements from the query to be simplified, wherein the semantic elements include entities, conditions, and attributes. The large language model replaces the specific values ​​corresponding to the entities, conditions, and attributes in the query to be simplified with corresponding entity placeholders ([Entity]), condition placeholders ([Condition]), and attribute placeholders ([Attribute]) to generate a structured logical representation.

[0069] In some exemplary embodiments, the prompt provided to the large language model can also instruct it to analyze and determine the desired final result type (i.e., the 'return type') of the user's query, for example, determining whether the user desires a list, a numerical statistical result, or a Boolean result. Predicting the return type helps facilitate the selection of subsequent operators, further improving the user experience and the targeted nature of the analysis.

[0070] For example, Figure 4 The query to be simplified is shown as "Among questions with over 500 views, which sport involving a ball has the highest ratio of number of injury-related questions to number of training-related questions?". The prompt provided to the large language model can be: "Please parse the following question to extract the entities, conditions, attributes, and the return type."

[0071] For example, Figure 4 In the parsing result shown, "questions" and "sport" in the original query are identified as entities ([Entity]), and "with over 500 views", "involving a ball", etc. are identified as conditions ([Condition]). Figure 4It vividly demonstrates the process of semantic parsing.

[0072] Step 3: Match the logical representation of the query to be simplified with the logical representations of a plurality of predefined standard semantic operators based on semantic similarity, and select candidate operators, each of which has at least one logical representation.

[0073] According to the present application, a series of standardized semantic operators can be predefined. These operators may include common data operations (such as filtering and aggregation) in traditional databases, and may also include semantic-level operations on unstructured data. In some embodiments, the predefined standard semantic operators may include some or all of the following: retrieval, filtering, aggregation, validation, sorting, comparison, refinement, translation, transformation, scanning, extraction, conceptualization, generation, explanation, grouping, clustering, classification, linking, set operation, integration, calculation, and summary. The following table is a description of multiple standard semantic operators predefined according to an exemplary embodiment of the present application.

[0074] Table 1 Predefined standard semantic operators and their descriptions

[0075]

[0076]

[0077] Each standard semantic operator has at least one logical representation. The logical representation is the embodiment of the core function and matching mode of the operator. In some embodiments, the logical representation of the semantic operator is a structured template, which is obtained by replacing the specific values ​​of semantic elements that generally represent a class of natural language fragments with similar logical structures with corresponding placeholders. The placeholders can be entity placeholders, condition placeholders, attribute placeholders, etc. The logical representation is structured and can be used to clearly represent the meaning of these natural language fragments at the logical level. For example, a Filter operator may have a logical representation "[Entity]that[Condition]", which can match natural language fragments such as "Movies that were made in the 2000s". A standard semantic operator can have multiple logical representations to recognize and process users' diverse natural language expressions. For example, the semantic operator Filter can be defined with multiple different logical expressions, such as "[Entity]that[Condition]", "[Entity]having[Condition]", or "[Entity]thatsatisfies[Condition]", so as to accurately identify the user's essentially the same filtering intention expressed in different wordings.

[0078] The definition of a standard semantic operator may also include input and output definitions and execution procedures.

[0079] The input and output definitions of standard semantic operators specify what input data or parameters are required when the operator is executed (for example, intermediate result data of a specific type, constant values, etc.), and what output results will be produced after execution (for example, new intermediate result data, Boolean values, text, etc.).

[0080] The execution process is used to define how the semantic operator completes its function. Each standard semantic operator can be executed through the corresponding physical operator. According to this application, there are two main ways to implement physical operators:

[0081] Method 1: Physical operators implemented through pre-programmed algorithms are similar to operators in traditional databases. They are executed using deterministic algorithms and are highly efficient. They are suitable for structured operations or simple semantic judgments.

[0082] The second method is to use large language models to process physical operators implemented by semantic operations. It is used to handle tasks that require deep semantic understanding, reasoning, generation, or complex judgment, such as filtering documents in the sports field based on whether the document content involves "teamwork."

[0083] Among the various standard semantic operators, some or all operators (eg, most of the semantic operators in Table 1) may have both types of physical operator implementations to enhance the flexibility of the system in processing tasks of different complexities.

[0084] The constructed set of semantic analysis operators provides the basis for query understanding and plan generation.

[0085] After obtaining the structured logical representation of the query to be simplified in step 2, step 3 can be used to find out which predefined standard semantic operators may be applicable to solving the current query or a part thereof.

[0086] In some implementations, screening candidate operators may include:

[0087] Convert the logical representation of the query to be simplified into a query semantic embedding vector. This process can be achieved through a semantic embedding model. The high-dimensional vector obtained by the conversion can capture the deep semantic meaning of the logical representation.

[0088] Using the query semantic embedding vector as the query vector, a nearest neighbor search is performed in a pre-built vector index to calculate the similarity between the query semantic embedding vector and the operator semantic embedding vector. The vector index stores the operator semantic embedding vector converted from the logical representation of the standard semantic operator. In some examples, cosine similarity or dot product similarity can be used to measure the similarity between the two semantic embedding vectors.

[0089] The standard semantic operators corresponding to the semantic embedding vectors of multiple operators with the highest similarity are selected as candidate operators. For example, by selecting multiple candidate operators instead of a single candidate operator, more choices and flexibility can be provided for subsequent re-ranking and optimization steps.

[0090] Figure 5 This example illustrates the operator matching process. The semantic parsing result is a query with placeholders. The system compares its embedding vector with the embedding vectors of pre-stored operator logical representations (e.g., "How many [Entity] are [Condition]," "[Entity] satisfies [Condition]," "[Entity] that [Condition]," etc.), selects the operator with the highest similarity as the candidate operator, and sorts them by similarity.

[0091] Step 4: Evaluate the degree to which the candidate operators solve the query to be simplified and their executability, and reorder the candidate operators.

[0092] Through a more detailed evaluation and re-ranking process, the candidate operators are further screened so that the screened operators are more suitable for simplifying the current query.

[0093] In some embodiments, evaluating and reordering may include:

[0094] Use a large language model to check whether the preconditions of each candidate operator have been met;

[0095] Using a large language model, evaluate the degree to which each candidate operator that meets the preconditions solves the query to be simplified, which can be completely solved, partially solved, or not solved at all;

[0096] The candidate operators are ranked based on the degree to which they solve the query to be simplified and the semantic similarity between the logical representation of the candidate operators and the logical representation of the query to be simplified.

[0097] A large language model can be used to check whether the preconditions of each candidate operator are met. An operator's preconditions typically refer to the availability of the input data required for its execution. The system maintains a dynamic set of available variables, typically intermediate results generated by the execution of other operators. The large language model determines whether the input variables required by the current candidate operator are included in this set. Only operators whose preconditions are met are considered executable.

[0098] For candidate operators that meet the preconditions, a large language model is used to evaluate the extent to which each operator can solve or partially solve the query to be simplified. The large language model makes judgments based on preset prompts, and the output may include fully solving, partially solving, or not solving.

[0099] Finally, according to the present application, the candidate operators can be comprehensively ranked based on the above evaluation results and in combination with semantic similarity. The most important ranking basis is the degree to which the candidate operator solves the query to be simplified. The higher the degree of solution, the higher the ranking. At the same time, the semantic similarity between the logical representation of the candidate operator and the logical representation of the query to be simplified can also be referenced. For example, similarity can be used to break the ranking between operators of the same degree. By re-ranking, an ordered list of operators can be obtained that is more likely to efficiently and successfully simplify the query.

[0100] Figure 6 The reordering process is shown as an example.

[0101] Step 5: Apply the candidate operators to simplify the query to be simplified in sequence according to the sorting results until the current query to be simplified is successfully simplified to obtain a simplified query.

[0102] Step 5 is the core step of query plan generation. It attempts to simplify the current query by applying the reordered candidate operators in sequence, thereby finding and applying an operator that can effectively promote query decomposition and obtain a simplified query that is closer to the final solution state.

[0103] In some embodiments, this may include:

[0104] Using a large language model, the expected outputs of the re-ranked candidate operators are sequentially applied to replace the matching fragments of the candidate operators in the query to be simplified, and the query after replacement is judged to see whether it meets the preset criteria.

[0105] The candidate operator of the query that meets the preset criteria obtained after the first replacement is regarded as the successfully applied operator, and the corresponding query that meets the preset criteria is regarded as the simplified query;

[0106] Generate intermediate result data and corresponding natural language descriptions representing the execution results of successfully applied operators, and add them to the set of available variables for use by subsequent operators.

[0107] According to this embodiment, the expected outputs of these candidate operators can be applied in sequence to replace the fragments of the query to be simplified that can be matched by the candidate operator, strictly following the order of the re-sorted candidate operator list, starting with the highest priority operator, using a large language model and a few sample example prompt words. After each replacement attempt, it can be determined whether the replaced query meets the preset criteria. Those skilled in the art can set the preset criteria as needed, for example, whether the generated query is well-structured, whether it correctly reflects the operation of the operator semantically, whether it is simpler or closer to a solvable state than the original query, whether the large language model completes the conversion with high confidence, etc.

[0108] The above attempt process will continue until the query obtained after substitution meets the preset criteria. The corresponding candidate operator is considered to be the operator successfully applied in this round of iteration, and the query obtained that meets the preset criteria is the simplified query in this round of iteration. Figure 7 The application process is schematically illustrated: using a large language model and prompts, a sequentially selected operator (such as a filter) and its expected output are used to rewrite (i.e., simplify) the query and obtain the simplified query.

[0109] Once the successfully applied operator and simplified query are determined, intermediate result data representing the execution result of the operator can be generated, and the corresponding natural language description can be generated using a large language model and added to the set of available variables for use by subsequent operators, for example, Figure 7The simplified newly added intermediate result data in is "questions withover 500views", and its data type is text document.

[0110] The above simplification process can be performed recursively. According to some embodiments, after obtaining the simplified query, a large language model can be used to determine whether the simplified query is a minimum semantic unit that cannot be decomposed any further. The minimum semantic unit includes a single entity, a numerical value, or a Boolean condition. For example, Figure 9 The final simplification is to "Group Name", which can be treated as a single entity.

[0111] If it is determined that the current simplified query is not the smallest semantic unit, it means that the query needs to be further decomposed and simplified. At this point, the process can return to Figure 1 In step 2, the current simplified query is used as the new query to be simplified, and semantic parsing is performed again (step 2), matching based on semantic similarity (step 3), candidate operators are reordered (step 4), and the query to be simplified is simplified (step 5). The recursive process ends when it is determined that the current simplified query is the smallest semantic unit that cannot be decomposed any further. In other words, the query can be considered to have been completely decomposed.

[0112] According to this embodiment, through recursive simplification, a complex initial natural language query can be gradually decomposed and transformed layer by layer until a set of indivisible semantic units forming the basis of an execution plan is obtained.

[0113] Step 6: When it is determined that the currently simplified query is a minimum semantic unit that cannot be decomposed any further, an execution plan is constructed according to the query simplification process, and the execution plan is executed using a parallel execution strategy based on topological sorting, and an analysis result is output.

[0114] During query simplification, each time an operator is successfully applied, its input (which may come from the original data or previous intermediate results) and output (new intermediate results) are recorded, forming a node in the plan. Dependencies between these nodes are also naturally formed. After the entire recursive simplification process is completed, the recorded sequence of operator applications and their dependencies constitute a complete, executable logical query plan.

[0115] A parallel execution strategy based on topological sorting can be used to execute the plan. For example, topological sorting can be performed based on the dependencies between operators in the execution plan. Operators whose preconditions are met and have no direct dependencies can be scheduled to different execution units for parallel execution.

[0116] During execution, you can also verify the intermediate results produced by the operator to determine whether they meet expectations. If the intermediate results do not meet expectations, you can analyze the cause based on the degree of deviation and adopt different adjustment strategies, such as adjusting the current execution plan, regenerating the execution plan, or rewriting the original natural language query to re-identify relevant data.

[0117] When all operators in the execution plan are successfully executed, the system will obtain the final analysis results and output them to the user.

[0118] In addition, the method of the present application may also include semantic parsing and preprocessing of the unstructured data involved before or during the execution of the operator in the execution plan, which can be performed in advance or when the operator needs to access the original data. The preprocessing step may include:

[0119] Perform word segmentation, entity recognition, and semantic annotation on the input unstructured data (such as text, documents, etc.) to generate an initial semantic feature vector;

[0120] Utilize a large language model to perform contextual analysis on these initial semantic feature vectors to obtain enhanced, context-sensitive semantic representations;

[0121] These enhanced context-sensitive semantic representations are then stored in a machine-readable structured intermediate format, such as a structured JSON or XML document, so that subsequent analysis operators can efficiently query and use them. This data preprocessing step helps improve the accuracy and efficiency of subsequent semantic analysis.

[0122] To further illustrate the technical solution of the present application, exemplary embodiments according to the present application are introduced below with reference to the accompanying drawings.

[0123] Figure 2 A system architecture diagram is generated for a plan according to an exemplary embodiment of the present application. Figure 2 The core architecture of the system, which begins with a user input query, is processed by the plan generation module, and ultimately results in an execution plan. Within the plan generation module, the system identifies relevant data for the query (for example, through index lookups) and utilizes a predefined set of standard semantic operators (including extraction, grouping, filtering, and comparison). The system then repeatedly performs a recursive simplification process, including operator matching, operator reordering, and query simplification, until the query is fully simplified, thereby constructing a complete execution plan.

[0124] Figure 3 This is an example diagram of an automated plan generation method according to an exemplary embodiment of the present application. Figure 3Taking the natural language query "Among questions with over 500 views, which sport involving a ball has the highest ratio of injury-related questions to training-related questions?" as an example, the main steps in its plan generation are demonstrated. First, the query is semantically parsed, converting it into a logical representation where specific values ​​are replaced by corresponding placeholders. Operator matching is then performed to obtain logical representations of matching candidate operators, such as "How many [Entity] are [Condition]" and "[Entity] satisfies [Condition]." Operator reordering is then performed, and candidate operators are applied sequentially based on the ranking results to simplify the query, for example, simplifying "Among questions with over 500 views" to "Among questions." The simplification process is recursive. After each simplification, if the query is not fully decomposed (i.e., the simplified query is not yet a minimal, irredecomposable semantic unit), the current simplified query is used as the new query to be simplified. The semantic parsing, operator matching, operator reordering, and query simplification steps are repeated until the query is fully decomposed. Then, based on the query simplification process, a logical plan, namely the execution plan, is formed.

[0125] Figure 4 Figure 2 is a schematic diagram illustrating semantic parsing of a natural language query according to an exemplary embodiment of the present invention. The natural language query to be parsed is "Among questions with over 500 views, which sport involving a ball has the highest ratio of injury-related questions to training-related questions?" The query is parsed using the large model by providing the prompt "Please parse the following question to extract the entities, conditions, attributes, and thereturn type." Figure 4The rightmost image shows the parsing result, where the identified entity semantic elements are marked in red, and the identified condition semantic elements are marked in green. The specific values ​​of these identified entities and conditions can then be replaced by the corresponding placeholders [Entity] and [Condition] to generate the logical representation of the query.

[0126] Figure 5 FIG. 1 is a schematic diagram of matching a suitable operator for a query based on semantic matching according to an exemplary embodiment of the present application. Figure 5 As shown in Figure 2, after obtaining the semantic parsing results, the logical representation of the query can be converted into an embedding vector for the query, namely the query semantic embedding vector. The similarity between this vector and the embedding vector of the logical representation of the pre-calculated and stored operator (i.e., the operator semantic embedding vector) is then calculated. Based on the calculated similarity values, multiple standard semantic operators with the highest similarity are selected as candidate operators. Figure 5 In the example shown, according to the similarity values, the candidate operators selected are "How many [Entity] are [Condition]", "[Entity] satisfy [Condition]", "[Entity] that [Condition]", etc., and their corresponding similarity values ​​are 0.83, 0.77, and 0.65, respectively.

[0127] Figure 6 Schematic diagram of reordering candidate operators according to an exemplary embodiment of the present application. Figure 6The left side of the figure shows a set of candidate operators, including "How many [Entity] are [Condition]," "[Entity] satisfies [Condition]," and "[Entity] that [Condition]." The first step in reranking is to check the preconditions of the candidate operators. This can be done by referencing existing intermediate results to verify whether the preconditions are satisfied. Specifically, this means that all required inputs for a candidate operator have been added to the current set of available variables. Next, the large language model is prompted with the following: "Please check whether the operator can solve any part of the query. If so, output the degree of solution (fully solving, partially solving), otherwise output not solving")" to evaluate the degree to which the operator solves the query. The large language model can output a result that is either fully solving, partially solving, or not solving. Finally, the candidate operators are reranked based on the degree to which each candidate operator solves the query to be simplified, combined with the semantic similarity between the candidate operator and the query to be simplified. Figure 6 In the example shown, the candidate operator with the second highest semantic similarity is ranked first after reordering.

[0128] Figure 7This figure shows a schematic diagram of query simplification using candidate operators according to an exemplary embodiment of the present application. The query to be simplified is "Among questions with over 500 views, which sport involving a ball has the highest ratio of injury-related questions to training-related questions?" First, the Filter operator is applied, with the logical representation of "[Entity] satisfies [Condition]" and the expected output is [Entity]. A prompt containing the above information, "Given the query [Query] and a matched logical representation (LR) of operator [OP] with expected output (OUTPUT), please rewrite the query with the operator by reducing the matched segment of the logical representation.", is provided to the large language model. The large language model successfully applies the operator to simplify the query, simplifying "Among questions with over 500 views" to "Among questions". At the same time, it also generates intermediate result data representing the execution result of the successfully applied operator Filter, that is, the intermediate result data "questions with over 500 views" with a text document data type, which will be added to the available variable set for use by subsequent related operators, such as as input or as a basis for judging whether the preconditions are met.

[0129] Figure 8 2 is a schematic diagram of determining whether the current simplified query is a minimum semantic unit that cannot be decomposed further according to an exemplary embodiment of the present application. Figure 8In the example shown, the prompt provided to the large model is "Check whether the initial query [Original query] has been fully resolved given the [Variable set description] and the current reduced query [Reduced query].". In this judgment, the original query is "Among questions with over 500 views, which sport involving a ball has the highest ratio of number of injury-related questions to number of training-related questions?", the reduced query is "Among questions, which sport involving a ball has the highest ratio of number of injury-related questions to number of training-related questions?", and the variable set description is the description of the currently available variable set. Figure 8 In the example shown, the large model determines that the currently simplified query needs to be further decomposed. Therefore, the currently simplified query "Among questions, which sport involving a ball has the highest ratio of number of injury-related questions to number of training-related questions?" will be used as the original query for the next round of simplification. After semantic parsing, operator matching, operator reordering, and query simplification, the large model will determine that the query has been completely decomposed. In other words, the currently simplified query has become the smallest semantic unit that cannot be decomposed any further.

[0130] Figure 9 A schematic diagram of complete query decomposition and plan generation according to an exemplary embodiment of the present application. Figure 9The process of gradually applying a series of standard semantic operators to decompose the natural language query "Among questions with over 500 views, which sport involving a ball has the highest ratio of number of injury-related questions to number of training-related questions?" is shown in a tree-like process (or directed acyclic graph structure).

[0131] Figure 9 The process shown mainly includes the following steps:

[0132] Step 1: Apply the Filter operator to the original query "Among questions with over 500 views, which sport involving a ball has the highest ratio of number of injury-related questions to number of training-related questions?" and perform preliminary filtering on the questions using the condition "with over 500 views".

[0133] Step 2: Apply the Extract operator to the results filtered in the previous step to extract information related to "sport".

[0134] Step 3: Apply the GroupBy operator to group the data by "sport".

[0135] Step 4: Apply the Filter operator to the grouped data and use the condition "about sports involving a ball" to further filter out groups related to ball sports.

[0136] Step 5: At this point, the query needs to find which group has the highest ratio of "injury-related" and "training-related" questions. To this end, these two types of questions will be processed separately in parallel.

[0137] Step 6 (left branch): Apply the Filter operator to filter out "injury-related" issues.

[0138] Step 7 (left branch): Apply the Count operator to the "injury-related" questions filtered out in step 6 to obtain the count result, which is recorded as "number of questions1".

[0139] Step ⑧ (right branch): Similar to step ⑥, apply the Filter operator to filter out "training-related" problems.

[0140] Step 9 (right branch): Apply the Count operator to the “training-related” questions selected in step 8 to obtain the count result, which is recorded as “number of questions2”.

[0141] Step ⑩: Take the two counting results ("number of questions1" and "number of questions2") obtained in steps ⑦ and ⑨ as input, apply the Compute operator, and calculate the ratio between them ("ratio of number1 to number2").

[0142] step Apply the Max operator to the proportion results of all groups calculated in step ⑩ to find the group with the highest proportion value (highest value) and finally output the name of the group "Group Name".

[0143] The figure clearly shows the complete process of recursively simplifying a complex analysis task by combining and applying different semantic operators, breaking it down into a series of simpler, executable subtasks, and ultimately obtaining the desired analysis results.

[0144] Figure 10 The following is a flow chart illustrating the execution of a plan according to an exemplary embodiment of the present application. First, the generated execution plan is executed from the bottom up according to the topology. During the execution process, intermediate result data is generated. A large model can be used to determine whether these intermediate result data meet expectations, and different levels of processing strategies can be adopted based on the judgment results:

[0145] Continue execution: If the intermediate result data basically meets expectations, you can continue execution;

[0146] Adjustment plan: As shown in the figure, the adjustment is minor and is performed based on the current execution plan. For example, some operators, parameters, or connection relationships in the current execution plan are adjusted.

[0147] Replanning: When a major flaw is discovered in the current plan, but the problem lies primarily in the plan logic itself rather than the initial data identification, deeper adjustments can be made, such as regenerating part or all of the execution plan.

[0148] Re-identify relevant data: This strategy can be enabled when the judgment problem may be caused by inaccurate or incomplete input data that was initially identified. For example, the original natural language query can be rewritten into a semantically equivalent but differently expressed form, and the modified query can be used to re-find and locate relevant data from the data source and re-plan based on the new data foundation.

[0149] The method for analyzing unstructured data provided in the embodiment of the present application significantly reduces the technical threshold for unstructured data analysis through technical means such as natural language query understanding, application of predefined semantic operators, automatic operator matching and reordering, and recursive query simplification, allowing users to gain in-depth data insights without complex programming. This method can automatically build and optimize the execution plan of complex queries, which not only greatly improves the analysis efficiency, but also ensures the accuracy of information extraction from massive unstructured data. In addition, the system also considers the monitoring and adjustment of the execution process at the execution level, further enhancing the reliability of the analysis results.

[0150] Figure 11 An electronic device provided for at least one embodiment of the present application includes a memory and a processor, wherein the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement the method for analyzing unstructured data described in any embodiment or implementation of the present application when executing the computer instructions.

[0151] At least one embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for analyzing unstructured data described in any embodiment or implementation of the present application.

[0152] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the data processing device embodiment is generally similar to the method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the method embodiment.

[0154] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0155] Embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier to be executed by a data processing device or to control the operation of the data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by the data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0156] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0157] Computers suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to such mass storage devices to receive data from them or to transmit data to them, or both. However, a computer does not necessarily have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0158] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0159] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of specific embodiments of specific inventions. Certain features described in multiple embodiments within this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claimed as such, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of the sub-combination.

[0160] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

[0161] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential sequence to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.

[0162] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included in the scope of protection of one or more embodiments of this specification.

Claims

1. A method for analyzing unstructured data, characterized in that include: receiving a natural language query input by a user as a query to be simplified; Use a large language model to perform semantic analysis on the query to be simplified and generate a logical representation of the query to be simplified; Matching the logical representation of the query to be simplified with the logical representations of a plurality of predefined standard semantic operators based on semantic similarity, screening out candidate operators, each of the standard semantic operators having at least one logical representation; Evaluate the degree and executability of the candidate operators in solving the query to be simplified, and re-rank the candidate operators; Apply the candidate operators to simplify the query in sequence according to the sorting results until the query is successfully simplified to obtain the simplified query. When it is determined that the currently simplified query is a minimum semantic unit that cannot be decomposed any further, an execution plan is constructed according to the query simplification process, and the execution plan is executed using a parallel execution strategy based on topological sorting, and an analysis result is output.

2. The method according to claim 1, characterized in that The multiple standard semantic operators include some or all of the following operators: Retrieve, Filter, Aggregate, Validate, OrderBy, Compare, Refine, Translate, Transform, Scan, Extract, Conceptualize, Generate, Explain, GroupBy, Cluster, Classify, Link, SetOP, Integrate, Compute, and Summarize.

3. The method according to claim 1, characterized in that The logical representation of the standard semantic operator is a structured template obtained by replacing the specific values ​​of the semantic elements in the natural language fragment with corresponding placeholders, which is used to represent the logical meaning of the natural language fragment. The placeholders include entity placeholders, condition placeholders and attribute placeholders.

4. The method according to claim 1, wherein Each of the standard semantic operators has at least one corresponding physical operator, which is implemented through a pre-programmed algorithm or implemented by processing semantic operations using a large language model. Some or all of the multiple standard semantic operators have both physical operators implemented through pre-programmed algorithms and physical operators implemented by processing semantic operations using a large language model.

5. The method according to claim 1, wherein A large language model is used to perform semantic analysis on the query to be simplified and generate a logical representation of the query to be simplified, including: Providing prompt words to the large language model to instruct the large language model to parse semantic elements from the query to be simplified, wherein the semantic elements include entities, conditions, and attributes; The large language model replaces the specific values ​​corresponding to entities, conditions, and attributes in the query to be simplified with corresponding entity placeholders, condition placeholders, and attribute placeholders to generate a structured logical representation.

6. The method according to claim 1, characterized in that Based on the semantic similarity between the logical representation of the query to be simplified and the logical representation of multiple predefined standard semantic operators, candidate operators are screened out, including: Convert the logical representation of the query to be simplified into a query semantic embedding vector; Using the query semantic embedding vector as the query vector, a nearest neighbor search is performed in a pre-built vector index to calculate the similarity between the query semantic embedding vector and the operator semantic embedding vector, wherein the vector index stores the operator semantic embedding vector converted from the logical representation of the standard semantic operator; The standard semantic operators corresponding to the semantic embedding vectors of multiple operators with the highest similarity are selected as candidate operators.

7. The method according to claim 1, characterized in that Evaluate the degree and executability of the candidate operators in solving the query to be simplified and re-rank the candidate operators, including: Use a large language model to check whether the preconditions of each candidate operator have been met; Using a large language model, evaluate the degree to which each candidate operator that meets the preconditions solves the query to be simplified, which can be completely solved, partially solved, or not solved at all; The candidate operators are ranked based on the degree to which they solve the query to be simplified and the semantic similarity between the logical representation of the candidate operators and the logical representation of the query to be simplified.

8. The method according to claim 1, characterized in that Apply the candidate operators to the query to be simplified in sequence according to the sorting results until the query to be simplified is successfully simplified. The simplified query is obtained, including: Using a large language model, the expected outputs of the re-ranked candidate operators are sequentially applied to replace the matching fragments of the candidate operators in the query to be simplified, and the query after replacement is judged to see whether it meets the preset criteria. The candidate operator of the query that meets the preset criteria obtained after the first replacement is regarded as the successfully applied operator, and the corresponding query that meets the preset criteria is regarded as the simplified query; Generate intermediate result data representing the execution results of the successfully applied operator and the corresponding natural language description, and add them to the available variable set for use by subsequent operators.

9. The method according to claim 1, characterized in that The method further includes, after obtaining the simplified query, determining whether the simplified query is a minimum semantic unit that cannot be further decomposed by: Using a large language model, it is determined whether the simplified query is a minimal semantic unit that cannot be decomposed any further. The minimal semantic unit includes a single entity, a numerical value, or a Boolean condition.

10. The method according to claim 1, characterized in that The method further comprises: When it is determined that the current simplified query is not the smallest semantic unit that cannot be decomposed any further, the current simplified query is used as a new query to be simplified, and semantic parsing, matching based on semantic similarity, candidate operators are reordered, and the query to be simplified is simplified again until it is determined that the current simplified query is the smallest semantic unit that cannot be decomposed any further.

11. The method according to claim 1, wherein The execution plan is executed using a parallel execution strategy based on topological sorting, including: Topological sorting is performed based on the dependencies between operators, and operators whose preconditions are met and have no dependencies on each other are executed in parallel.

12. The method according to claim 1, characterized in that The method also includes, during the execution of the execution plan, if the intermediate result data generated does not meet expectations, adjusting the current execution plan, regenerating the execution plan, or rewriting the original natural language query to re-identify relevant data.

13. The method according to claim 1, wherein The method further includes the steps of semantically parsing and preprocessing the unstructured data, including: Perform word segmentation, entity recognition, and semantic annotation on the input unstructured data to generate an initial semantic feature vector; Utilize a large language model to perform contextual analysis on the initial semantic feature vector to obtain an enhanced context-dependent semantic representation. The enhanced context-sensitive semantic representation is stored in a structured machine-readable intermediate format.

14. An electronic device, characterized in that: The device includes a memory and a processor, wherein the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement the method according to any one of claims 1 to 13 when executing the computer instructions.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Graph compiling method and device for large language model fusion operator and storage medium

    CN118034660A

  • Index acquisition method and device, electronic equipment and computer readable storage medium

    CN118673038A

  • Complex query-oriented automatic interactive large language model pipeline arrangement method, system and equipment and storage medium

    CN120030048A

  • Method and System for Sentiment Analysis of News Articles based on AI

    KR102597357B1

  • Computer implemented methods for the automated analysis or use of data, including use of a large language model

    WO2023161630A1