Ontology-based data set validity confirmation method

Through the ontology-based data set validity confirmation method, using feature extractor, domain ontology and semantic reasoning machine, the problems of inaccurate semantic understanding and difficulty in standardization are solved, the effectiveness confirmation of data sets is realized, and the efficiency and quality of AI software development are improved.

CN120354856APending Publication Date: 2025-07-22SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510411061.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing semantic understanding solutions generally have problems such as inaccurate understanding and difficulty in standardizing.

Method used

The ontology-based data set validity confirmation method is used to extract the semantic information of the target data set through a pre-trained feature extractor, and normalize and deep semantic inference are used to perform normalization and deep semantic inference. Combined with software engineering requirements analysis, SPARQL query is generated to filter out the conditions to confirm the validity of the data set.

Benefits of technology

The validity confirmation of the data set before training is achieved, which improves the efficiency and quality of AI software development, reduces development costs and improves the performance of the final product.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354856A_ABST
    Figure CN120354856A_ABST
Patent Text Reader

Abstract

The invention discloses an ontology-based data set validity confirmation method. The method comprises the following steps of obtaining a to-be-evaluated target data set; semantic information in the target data set is extracted through a pre-trained feature extractor; normalizing the extracted semantic information concept through the domain ontology, and converting the extracted semantic information concept into a standardized semantic representation form; deducing deep semantics of semantic information through a semantic inference engine to obtain RDF format data; necessary conditions of a target function are analyzed and extracted based on software engineering requirements, and a clear standard is provided for subsequent data set evaluation; sPARQL query is generated according to the domain ontology, images meeting any query condition are screened out, and the validity of the target data set is confirmed. According to the ontology-based data set validity confirmation method provided by the invention, scene graph and ontology mapping are combined, the accuracy of semantic understanding is improved, standardized expression of semantics is realized, flexible demand matching is supported, and the method has good expansibility and maintainability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and knowledge graphs, and particularly to a method for validating the effectiveness of a dataset based on ontology. Background Art

[0002] The existing neural network model structures tend to be stable, and the research focus has shifted to data. The quality of training data determines the performance of the AI system. The decision-making model consists of two parts: feature extraction and decision generation. Among them, the feature extraction process is the premise of decision-making, and the upper bound of the application program function can be estimated through the performance of the SOTA-level model in this field.

[0003] Semantic understanding is the basis for evaluating the effectiveness of a dataset. Only after correctly and fully understanding the data content can a reliable evaluation and judgment be made.

[0004] Currently, the following several technical solutions mainly exist for semantic understanding:

[0005] (1) Rule-based method: Extract semantic information by manually defining grammar rules and patterns, but this method has poor scalability and requires a large amount of manual maintenance;

[0006] (2) Deep learning-based method: Use a neural network to directly learn semantic representations from text. This method lacks interpretability and it is difficult to ensure semantic accuracy;

[0007] (3) Knowledge graph-based method: Perform semantic matching using a predefined knowledge base. This method relies on a high-quality knowledge base architecture and has a high cost.

[0008] Therefore, the existing semantic understanding solutions generally have problems of inaccurate understanding and difficulty in standardization. Summary of the Invention

[0009] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is that the existing semantic understanding solutions generally have problems of inaccurate understanding and difficulty in standardization. Therefore, the present invention provides a method for validating the effectiveness of a dataset based on ontology, which combines a scene graph and ontology mapping to improve the accuracy of semantic understanding, achieve standardized expression of semantics, support flexible requirement matching, and has good scalability and maintainability.

[0010] To achieve the above object, the present invention provides a method for validating the effectiveness of a dataset based on ontology, including the following steps:

[0011] Obtain a target dataset to be evaluated;

[0012] Extract semantic information in the target dataset through a pre-trained feature extractor;

[0013] Normalize the semantic information concepts extracted through the domain ontology and convert them into a standardized semantic representation form;

[0014] Infer the deep semantics of the semantic information through a semantic inference engine to obtain data in RDF format;

[0015] Extract the necessary conditions of the target function based on the software engineering requirements analysis to provide a clear standard for subsequent dataset evaluation;

[0016] Generate a SPARQL query according to the domain ontology, screen out the images that meet any query condition, and confirm the validity of the target dataset.

[0017] Furthermore, the target dataset includes a large number of images, and each image represents a different scene or situation.

[0018] Furthermore, extract the semantic information in the target dataset through a pre-trained feature extractor. Specifically, use the Faster R-CNN model as the pre-trained feature extractor to extract the semantic information of each image in the target dataset in the way of scene graph generation / detection, and the output results include entities and entity-relationship-entity triples.

[0019] Furthermore, the processing flow of the Faster R-CNN model includes the following steps:

[0020] Apply a convolutional neural network to the input images in the target dataset to extract feature maps, including scene graphs;

[0021] Generate candidate regions through a region proposal network, and screen out candidate regions with low overlap and high scores through the non-maximum suppression algorithm;

[0022] Apply RoI pooling to each candidate region to map candidate regions of different sizes into a fixed-size feature representation;

[0023] Perform object classification and bounding box regression on the pooled features to determine the object category and precise location;

[0024] Analyze the spatial and semantic relationships between objects to generate entity-relationship-entity triples;

[0025] Integrate the object and the relationship information between them to form a complete scene graph.

[0026] Furthermore, normalize the semantic information concepts extracted through the domain ontology and convert them into a standardized semantic representation form, specifically including normalizing the semantic information output by the feature extractor into a standardized semantic representation form using the domain ontology.

[0027] Furthermore, the normalization of semantic information concepts includes the following steps:

[0028] Map the entities recognized by the feature extractor to the standard concepts defined in the domain ontology;

[0029] Map the relationships recognized by the feature extractor to the standard relationships defined in the domain ontology;

[0030] Standardize the various attributes of the entities recognized by the feature extractor and map them to the standard attributes defined in the domain ontology;

[0031] Process the uncertain information output by the feature extractor and finally convert the original semantic information into a standardized semantic representation that conforms to the ontology.

[0032] Furthermore, infer the deep semantics of the semantic information through a semantic inference engine to obtain data in RDF format, specifically including using the HermiT inference engine to infer deep semantic relationships based on the standardized semantic concepts and the inference rules defined in the domain ontology, and generating semantic data in RDF format.

[0033] Furthermore, the RDF format is a triple format, including subject - predicate - object; the specific method for generating semantic data in RDF format includes entity representation, relationship representation, attribute representation, and confidence representation.

[0034] Furthermore, extract the necessary conditions for the target function based on the software engineering requirements analysis, providing a clear standard for the subsequent dataset evaluation. This is achieved through the software engineering requirements analysis method to determine the key entities, relationships, and scenario states that need to be recognized to achieve the target function, forming a list of functional requirements, specifically including:

[0035] Collect the requirement information related to the target function, analyze the obtained requirement information, and decompose it into more specific and operable functional requirements;

[0036] Based on the results of the requirements analysis, extract the necessary conditions for achieving the target function, including the entity types, relationship types, and scenario states that the system needs to recognize;

[0037] Classify the extracted necessary conditions by priority, distinguishing core necessary conditions and secondary conditions;

[0038] Organize the extracted necessary conditions into a structured list of functional requirements, clarifying the description, priority, and acceptance standard information for each requirement; this list will be used as the standard for evaluating the dataset.

[0039] Furthermore, generate SPARQL queries according to the domain ontology, screen out the images that meet any query condition, and confirm the validity of the target dataset. Specifically, generate SPARQL queries according to the list of functional requirements and the domain ontology, retrieve the images that meet the requirements, and evaluate the validity of the dataset based on the number and coverage of the images that meet the conditions.

[0040] Technical effect

[0041] A method for validating the effectiveness of a dataset based on ontology provided by the present invention realizes validating the effectiveness of the dataset before training, can timely screen out data that meets the requirements, avoid ineffective training, greatly improve the efficiency and quality of AI software development, thereby reducing the development cost and enhancing the performance of the final product.

[0042] The following will further illustrate the concept, specific structure and technical effects generated by the present invention with reference to the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. Brief description of the drawings

[0043] Figure 1 is a schematic flowchart of a method for validating the effectiveness of a dataset based on ontology according to a preferred embodiment of the present invention;

[0044] Figure 2 is a scene effect diagram of a method for validating the effectiveness of a dataset based on ontology according to a preferred embodiment of the present invention;

[0045] Figure 3 is a schematic flowchart of a method for validating the effectiveness of a dataset based on ontology according to a preferred embodiment of the present invention;

[0046] Figure 4 is a scene effect diagram of a method for validating the effectiveness of a dataset based on ontology according to a preferred embodiment of the present invention. Detailed implementation manners

[0047] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] In the following description, specific details such as specific internal programs and technologies are put forward for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0049] With the development of artificial intelligence technology, more and more software systems are developed in an AI manner. During the development of these systems, the quality of the dataset directly affects the performance and functionality of the final software. However, in the existing technology, it is often only possible to judge whether a dataset is suitable for developing the target software after training, and this lagging judgment method leads to a waste of a large amount of computing resources and time. Judging whether a dataset contains the necessary elements for implementing software functions before training can effectively improve the efficiency and success rate of AI software development.

[0050] To solve this problem, in the embodiments of the present invention, a method for validating the effectiveness of a dataset based on ontology is provided. First, obtain the target dataset to be evaluated; then, extract the semantic information in the dataset through a pre-trained feature extractor; next, normalize the extracted semantic concepts through a domain ontology and infer the deep semantics; finally, based on the necessary conditions extracted from the software engineering requirements analysis, evaluate whether the dataset contains the elements required to achieve the target function. A method for validating the effectiveness of a dataset based on ontology in the present invention can, by validating the effectiveness of the dataset before training, promptly screen out the compliant data, avoid ineffective training, greatly improve the efficiency and quality of AI software development, thereby reducing the development cost and enhancing the performance of the final product.

[0051] As Figure 1-2 shown, the present invention provides a method for validating the effectiveness of a dataset based on ontology, including the following steps:

[0052] Step 100, obtain the target dataset to be evaluated; wherein, the target dataset includes a large number of images, and each image represents a different scene or situation.

[0053] Step 200, extract the semantic information in the target dataset through a pre-trained feature extractor; specifically, use the Faster R-CNN model as the pre-trained feature extractor, and extract the semantic information of each image in the target dataset in the way of scene graph generation / detection, and the output result includes entities and entity-relationship-entity triples;

[0054] The processing flow of the Faster R-CNN model includes the following steps:

[0055] Apply a convolutional neural network to the input image in the target dataset to extract a feature map, including a scene graph;

[0056] Generate candidate regions through a region proposal network, and screen out candidate regions with low overlap and high scores through a non-maximum suppression algorithm;

[0057] Apply RoI pooling to each candidate region to map candidate regions of different sizes into a fixed-size feature representation;

[0058] Perform object classification and bounding box regression on the pooled features to determine the object category and precise location;

[0059] Analyze the spatial and semantic relationships between objects to generate entity-relationship-entity triples;

[0060] Integrate object and relationship information between them to form a complete scene graph.

[0061] Step 300, normalize the semantic information concepts extracted through the domain ontology and convert them into a standardized semantic representation form; specifically, use the domain ontology to normalize the semantic information output by the feature extractor into a standardized semantic representation form.

[0062] The normalization of semantic information concepts includes the following steps:

[0063] Map the entities recognized by the feature extractor to the standard concepts defined in the domain ontology;

[0064] Map the relationships recognized by the feature extractor to the standard relationships defined in the domain ontology;

[0065] Standardize the various attributes of the entities recognized by the feature extractor and map them to the standard attributes defined in the domain ontology;

[0066] Process the uncertain information output by the feature extractor and finally convert the original semantic information into a standardized semantic representation that conforms to the ontology.

[0067] Step 400, infer the deep semantics of the semantic information through a semantic inference engine to obtain data in RDF format; specifically, use the HermiT inference engine to infer deep semantic relationships based on the normalized semantic concepts and the inference rules defined in the domain ontology, and generate semantic data in RDF format. The RDF format is a triple format, including subject-predicate-object; the specific method for generating semantic data in RDF format includes entity representation, relationship representation, attribute representation, and confidence representation.

[0068] Step 500, extract the necessary conditions for the target function based on software engineering requirements analysis to provide clear criteria for subsequent dataset evaluation; extracting the necessary conditions for the target function based on software engineering requirements analysis to provide clear criteria for subsequent dataset evaluation is to determine the key entities, relationships, and scenario states required to achieve the target function through the requirements analysis method of software engineering, and form a list of functional requirements, specifically including:

[0069] Collect requirement information related to the target function, analyze the obtained requirement information, and decompose it into more specific and operable functional requirements;

[0070] Based on the results of requirement analysis, extract the necessary conditions for achieving the target functions, including the entity types, relationship types, and scenario states that the system needs to identify;

[0071] Divide the extracted necessary conditions into priorities to distinguish core necessary conditions and secondary conditions;

[0072] Organize the extracted necessary conditions into a structured list of functional requirements, clarifying the description, priority, and acceptance criteria information for each requirement; this list will serve as the standard for evaluating the dataset.

[0073] Step 600, generate a SPARQL query based on the domain ontology, filter out the images that meet any query condition, and confirm the validity of the target dataset. Generate a SPARQL query based on the domain ontology, filter out the images that meet any query condition, and confirm the validity of the target dataset. Specifically, generate a SPARQL query based on the list of functional requirements and the domain ontology, retrieve the images that meet the requirements, and evaluate the validity of the dataset based on the number and coverage of the images that meet the conditions.

[0074] The present invention will illustrate a method for confirming the validity of a dataset based on ontology of the present invention with a specific example.

[0075] The scenario of the embodiment of the present application is applied to a scenario as shown in Figure 4 This scenario includes a dataset validity confirmation system. Among them, the dataset validity confirmation system includes a semantic extraction module, a semantic modeling module, a statistical verification module, a logical verification module, and a dataset evaluation module. When a user needs to develop an AI software system with a specific function, the dataset validity confirmation system obtains the target dataset to be evaluated and confirms the validity of the target dataset to be evaluated by adopting the implementation manner of the embodiment of the present application.

[0076] In the above application scenario, although the action description of the implementation manner of the present application is performed by the dataset validity confirmation system, however, the present application is not limited in terms of the execution subject, as long as the actions disclosed in the implementation manner of the present application are executed.

[0077] The above scenario is only a scenario example provided by the embodiment of the present application, and the embodiment of the present application is not limited to this scenario.

[0078] Next, in conjunction with the accompanying drawings, the specific implementation manners of the method and related devices for confirming the validity of the dataset in the embodiment of the present application will be described in detail through embodiments.

[0079] As shown in Figure 3 It shows a schematic flowchart of a method for confirming the validity of a dataset in an embodiment of the present application. In this embodiment, the method includes the following steps:

[0080] Step 100: Obtain the target dataset to be evaluated.

[0081] To confirm the validity of the dataset, it is first necessary to obtain the target dataset to be evaluated. Specifically, for example, in the scenario of autonomous driving system development, when a user needs to develop an autonomous driving system that can identify various objects in the traffic environment, the collected traffic scene image dataset can be used as the target dataset to be evaluated. The dataset validity confirmation system obtains this target dataset to be evaluated, and subsequently, it is necessary to confirm the validity of the target dataset to be evaluated to determine whether the dataset contains the necessary elements for implementing the autonomous driving function.

[0082] The target dataset usually contains a large number of images, and each image represents a different scene or situation. For example, in the development of an autonomous driving system, the target dataset may contain images of various traffic scenes, such as urban roads, highways, intersections, and road scenes under various weather conditions. These images need to cover all kinds of situations that the system will face to ensure that the trained AI model can adapt to the complex and changing actual environment.

[0083] Step 200: Extract the semantic information in the target dataset through a pre-trained feature extractor.

[0084] Traditional dataset evaluation methods usually rely on manual inspection or simple statistical analysis, and it is difficult to comprehensively evaluate whether the semantic information in the dataset meets the software function requirements. However, the embodiments of this application use a pre-trained feature extractor to automatically extract the semantic information in the dataset, which can more comprehensively and accurately understand the content of the dataset.

[0085] In the embodiments of this application, the pre-trained feature extractor, for example, uses the Faster R-CNN model, which has been pre-trained through the Visual Genome dataset and has semantic embedding through the GloVe word vector model. Faster R-CNN is an efficient object detection algorithm that can simultaneously identify objects in an image and their positions. Through pre-training on the Visual Genome dataset, this model has learned the ability to recognize various common objects and the relationships between them.

[0086] Use the pre-trained Faster R-CNN model as the feature extractor to extract the semantic information of each image in the target dataset in the way of scene graph generation / detection, and the output results include entities and (entity-relationship-entity) triples.

[0087] Specifically, the pre-trained Faster R-CNN model first performs region proposals on the image to identify regions that may contain objects, and then classifies and regresses the bounding boxes of these regions to determine the categories and precise positions of each object. Next, the model analyzes the spatial relationships and interactions between objects to generate (entity-relationship-entity) triples, forming a scene graph representation. For example, for an image containing "a pedestrian is crossing the road and a car is stopped in front of the traffic light", Faster R-CNN will extract entities such as "pedestrian", "road", "car", "traffic light", and triple relationships such as "pedestrian-cross-road", "car-stop-in-front-of-traffic-light".

[0088] More specifically, the processing flow of the Faster R-CNN model includes the following steps:

[0089] Step 201: Apply a convolutional neural network (CNN) to the input image to extract a feature map.

[0090] In this step, the input image first passes through a series of convolutional layers and pooling layers to generate a feature map containing rich feature information. These feature maps capture low-level features such as edges, textures, and shapes in the image, as well as high-level features such as object parts and complete objects. For the embodiments in this application, classic network architectures such as ResNet and VGG are usually used as the backbone for feature extraction.

[0091] Step 202: Generate candidate regions through a Region Proposal Network (RPN).

[0092] The RPN is a key component of Faster R-CNN. It slides a small window over the feature map and predicts multiple regions that may contain objects (called anchor boxes) at each position. The RPN predicts two values for each anchor box: a binary classification score indicating the likelihood that the region contains an object; and a bounding box regression value for adjusting the position of the anchor box to better fit the object. Through the non-maximum suppression (NMS) algorithm, candidate regions with low overlap and high scores are selected to reduce redundancy.

[0093] Step 203: Apply RoI (Region of Interest) pooling to each candidate region to extract a fixed-size feature representation.

[0094] Since the candidate regions are of different sizes, and the fully connected layer requires a fixed-size input, the RoI pooling operation is used to map candidate regions of different sizes into a fixed-size feature representation. This step ensures the input consistency for subsequent classification and regression operations.

[0095] Step 204: Perform object classification and bounding box regression on the pooled features to determine the object category and precise location.

[0096] This step further processes each candidate region through a fully connected layer, outputting two results: a multi-classification result indicating the category of the object in that region (such as pedestrian, vehicle, traffic light, etc.); a bounding box regression result for further precise positioning of the object. For example, in an autonomous driving scenario, the model may identify various entity categories such as "pedestrian", "vehicle", "road", "traffic signal", etc.

[0097] Step 205: Analyze the spatial and semantic relationships between objects to generate (entity - relation - entity) triples.

[0098] After identifying each object in the image, the model further analyzes the relationships between the objects. This is usually achieved through another neural network module that receives the visual and spatial features of object pairs as inputs and predicts the type of relationship between them. For example, triples relationships such as "pedestrian - crossing - road", "vehicle - stopping in front of - traffic light", "vehicle - on - road" may be identified. These relationship types are usually selected from a predefined set of relationships, such as the relationship categories defined in the Visual Genome dataset.

[0099] Step 206: Integrate the object and relationship information to form a complete scene graph representation.

[0100] Finally, integrate the identified objects and the relationships between them into a unified scene graph representation. A scene graph is a structured graph representation where nodes represent objects and edges represent the relationships between objects. This representation captures the semantic content of the image and facilitates subsequent semantic analysis and reasoning.

[0101] Through the above steps, the Faster R - CNN model can extract rich semantic information from the image, including entities and their categories, locations, attributes, and various relationships between entities. This information provides a basis for subsequent semantic normalization and reasoning.

[0102] Step 300: Normalize the semantic concepts in the semantic information through a domain ontology.

[0103] The semantic concepts output by the feature extractor are usually relatively rough and may have inconsistencies, redundancies, or ambiguities. To perform effective semantic analysis and reasoning, these semantic concepts need to be normalized. The embodiments of this application use a domain ontology to achieve this goal, ensuring that the semantic concepts conform to the standards and requirements of a specific domain.

[0104] Normalize the semantic concepts (including entities and relationships) output by the feature extractor into a standardized semantic representation using a domain ontology.

[0105] Specifically, a domain ontology is a formal knowledge representation that defines the concepts, attributes, and relationships in a specific domain. In the field of autonomous driving, the ontology may include concepts such as "vehicle", "pedestrian", "road", "traffic signal", etc., and various relationships between them such as "located at", "moving", "interacting", etc. The ontology also defines the hierarchical relationships between concepts, such as "sedan", "truck", "bus" are all subclasses of "vehicle".

[0106] The semantic concept normalization process may include the following steps, for example:

[0107] Step 301: Standardize entity concepts

[0108] In this step, map the entities recognized by the feature extractor to the standard concepts defined in the domain ontology. For example, the feature extractor recognizes objects in an image as "car", "automobile", "sedan", etc., all of which can be standardized to the "vehicle" concept in the ontology. This standardization process is usually achieved with the help of a thesaurus, word embedding similarity calculation, or specialized mapping rules.

[0109] More specifically, entity concept standardization can be achieved through the following methods:

[0110] (1) Direct mapping: Establish a mapping table for common entity concepts and directly map the concepts output by the feature extractor to the standard concepts in the domain ontology. For example, "car", "automobile", "vehicle" are all mapped to the standard concept "vehicle".

[0111] (2) Similarity calculation based on word embedding: Use word embedding models such as GloVe, Word2Vec, etc. to calculate the semantic similarity between the concepts output by the feature extractor and the standard concepts in the ontology, and select the standard concept with the highest similarity as the mapping result. For example, calculate the similarity between "sedan" and each concept in the ontology and find that the similarity with "vehicle" is the highest, so map "sedan" to "vehicle".

[0112] (3) Hierarchical mapping: According to the hierarchical relationships between concepts defined in the ontology, map specific concepts to the appropriate hierarchical levels. For example, the feature extractor may recognize "SUV", which is a subclass of "vehicle". Depending on the granularity of the analysis requirements, it can be retained as "SUV" or mapped to the more general "vehicle" concept.

[0113] Step 302: Standardize relationship concepts

[0114] Similarly, map the relationships recognized by the feature extractor to the standard relationships defined in the domain ontology. For example, spatial relationships such as "located at", "on...", "in..." may all be mapped to the standard relationship "spatiallyLocatedAt". This step ensures that the same relationships with different expressions can be uniformly represented for subsequent analysis.

[0115] Specific methods for standardizing relationship concepts include:

[0116] (1) Relationship type classification: Classify the relationships recognized by the feature extractor into several categories such as spatial relationships (e.g., "on...", "next to..."), action relationships (e.g., "driving", "crossing"), functional relationships (e.g., "used for", "control"), etc., and then standardize within each category.

[0117] (2) Relationship derivation rules: Define a series of rules to derive standard relationships based on the semantic characteristics of the relationships. For example, "The vehicle stops in front of the traffic light" can derive two standard relationships: "vehicle - spatiallyLocatedNear - traffic light" and "vehicle - respondTo - traffic light".

[0118] (3) Context - related mapping: Consider the types of entities at both ends of the relationship to determine the most suitable standard relationship. For example, although both "person - in - vehicle" and "vehicle - on - road" use "in" to represent the relationship, they should be mapped to different standard relationships, such as "isInsideOf" and "isOnTopOf".

[0119] Step 303: Attribute standardization

[0120] In addition to entities and relationships, the feature extractor may also recognize various attributes of entities, such as color, size, status, etc. These attributes also need to be standardized and mapped to the standard attributes defined in the domain ontology. For example, standardize "red", "scarlet", etc. to the attribute "color:red".

[0121] Specific methods for attribute standardization include:

[0122] (1) Attribute type recognition: First, recognize the type of the attribute, such as color, size, material, status, etc., and then standardize within each type.

[0123] (2) Attribute value normalization: For quantitative attributes, unify the units and ratios; for qualitative attributes, unify the expressions. For example, standardize "very large", "huge" to "size:large".

[0124] (3)Boolean conversion of attributes: In some cases, complex attributes can be converted into a series of Boolean attributes. For example, convert "the traffic light shows red" to "trafficLight.isRed: true".

[0125] Step 304: Uncertainty handling

[0126] The concepts output by the feature extractor sometimes carry uncertainties, such as confidence scores. During the normalization process, these uncertainty information needs to be properly handled. For example, a confidence threshold can be set to only retain concepts with high confidence; or the original confidence information can be retained for the normalized concepts for subsequent reasoning consideration.

[0127] Specific methods for uncertainty handling include:

[0128] (1) Confidence filtering: Set a confidence threshold (such as 0.7), and only normalize concepts that exceed the threshold, discarding recognition results with low confidence to reduce error propagation.

[0129] (2) Multiple candidate retention: When there are multiple possible mapping results for an entity or relationship, retain the several candidates with the highest confidence and consider multiple possibilities in subsequent reasoning.

[0130] (3) Confidence transfer: During the normalization process, adjust the confidence according to the certainty of the mapping. For example, direct mapping may maintain the original confidence, while fuzzy mapping based on similarity may reduce the confidence.

[0131] Through the above steps, the original semantic concepts output by the feature extractor are converted into a standardized semantic representation that conforms to the domain ontology, laying a foundation for subsequent semantic reasoning. This standardization process eliminates the inconsistency of expressions and improves the accuracy and reliability of semantic analysis.

[0132] Step 400: Infer the deep semantics of semantic information through a semantic reasoner to obtain data in RDF format.

[0133] It should be noted that although the normalized semantic concepts have a unified representation form, they may still only contain explicitly expressed semantic information. However, many important semantic information is implicit and needs to be obtained through reasoning. The embodiments of this application use a semantic reasoner to infer deep semantics and enrich the semantic representation of the dataset.

[0134] In an optional implementation manner of the embodiments of this application, step 400 is specifically: using the HermiT reasoner to infer deep semantic relationships based on the normalized semantic concepts and the inference rules defined in the domain ontology, and generating semantic data in RDF (Resource Description Framework) format.

[0135] Specifically, HermiT is a description logic-based inference engine that can perform various inference tasks based on ontologies and instance data, such as concept satisfiability checking, classification, instance retrieval, etc. In the embodiments of this application, the HermiT inference engine is mainly used to perform the following inference tasks:

[0136] Step 401: Type inference

[0137] Based on the attributes and relationships of known entities, infer more specific types of entities. For example, if an entity is identified as a "traffic signal" and has the attribute of "displaying a red signal", it can be inferred that it is a type of "traffic light". Type inference is usually based on the hierarchical relationships between types defined in the ontology and the necessary conditions of types.

[0138] More specifically, type inference can be achieved through the following methods:

[0139] (1) Inference based on attributes: Infer the type of an entity based on the attribute values of the entity. For example, if an entity has attributes such as "has wheels", "has an engine", and "can carry people", it can be inferred that it is of the type "vehicle".

[0140] (2) Inference based on relationships: Infer the type of an entity based on the relationships the entity participates in. For example, if an entity is the subject of the "driving" relationship, it can be inferred that it is a "person"; if it is the object of the "driving" relationship, it can be inferred that it is a "vehicle".

[0141] (3) Composite inference: Consider multiple information sources, such as the initial type, attributes, relationships, etc. of the entity, and jointly infer the most specific type of the entity. For example, an entity initially identified as a "vehicle", if it also has features such as "has a passenger cabin", "has multiple seats", and "has a public transportation function", can be further inferred as a "bus".

[0142] Step 402: Relationship inference

[0143] Based on the known entity relationships, infer implicit relationships. For example, if it is known that "Pedestrian A is on the sidewalk" and "next to the sidewalk is Road B", it can be inferred that "Pedestrian A is near Road B". Relationship inference is usually based on the relationship characteristics defined in the ontology, such as transitivity, symmetry, etc., and the inclusion relationships between relationships.

[0144] The specific methods of relationship inference include:

[0145] (1) Transitive relationship inference: If the relationship R is transitive, and there exists a-R-b and b-R-c, then it can be inferred that a-R-c. For example, if "A is inside B" and "B is inside C", then it can be inferred that "A is inside C".

[0146] (2) Inverse relationship inference: If relationship R has an inverse relationship R', and there exists a - R - b, then it can be inferred that b - R' - a. For example, if "A contains B", then it can be inferred that "B is contained in A".

[0147] (3) Relationship combination inference: Infer a new relationship based on the combination of two or more relationships. For example, if "a pedestrian is on the sidewalk" and "the sidewalk is next to a lane", then it can be inferred that "the pedestrian is near the lane", which is very important for safety analysis.

[0148] (4) Context - related inference: Consider the type and attributes of entities to infer relationships in a specific context. For example, if "Vehicle A is close to Vehicle B" and "the speed of Vehicle A is greater than that of Vehicle B", then it can be inferred that "Vehicle A may overtake Vehicle B", which is valuable for traffic flow analysis.

[0149] Step 403: Attribute inference

[0150] Based on the type, known attributes, and relationships of entities, infer the implicit attributes of entities. For example, if an entity is classified as an "emergency vehicle", it can be inferred that it has the attribute of "right of way", even if this attribute is not directly observed.

[0151] Specific methods of attribute inference include:

[0152] (1) Type - based attribute inheritance: According to the type hierarchy defined in the ontology, subtypes inherit the attributes of the supertype. For example, if the type "vehicle" has the attribute of "can move", then a "sedan", which is a subtype of "vehicle", also inherits this attribute.

[0153] (2) Relationship - based attribute inference: Infer the attributes of an entity from the relationships it participates in. For example, if "Vehicle A is on Lane B" and "Lane B is a one - way street", then it can be inferred that the "driving direction of Vehicle A" should conform to the one - way street regulations.

[0154] (3) Rule - based inference: Infer attributes according to predefined if - then rules. For example, the rule "If a vehicle approaches an intersection and the traffic signal is red, then the vehicle needs to stop" can infer the "needs to stop" attribute of the vehicle.

[0155] Step 404: Scenario state inference

[0156] Based on the comprehensive analysis of all entities, relationships, and attributes, infer the state of the entire scenario. For example, by analyzing vehicle positions, speeds, traffic signal states, etc., scenario states such as "smooth traffic at the intersection" or "potential collision risk" can be inferred.

[0157] Specific methods of scenario state inference include:

[0158] (1) Event recognition: Identify events based on the state changes of entities. For example, if the "position of Vehicle A" changes significantly in several consecutive frames, it can be inferred that "Vehicle A is moving".

[0159] (2) Risk assessment: Assess potential risks based on the spatial relationships and dynamic attributes between entities. For example, if "a pedestrian is crossing the road" and "a vehicle is approaching rapidly", it is inferred that "there is a risk of collision".

[0160] (3) Intention inference: Infer the intention of an entity based on its behavior pattern. For example, if "the vehicle decelerates" and "the turn signal is on", it can be inferred that "the vehicle is preparing to turn".

[0161] (4) Scene classification: Classify the entire scene based on the states and relationships of multiple entities. For example, classify the scene into states such as "congested" or "smooth" according to vehicle density, average speed, etc.

[0162] Step 405: RDF data generation

[0163] Represent all inference results in the standard RDF triple format (subject - predicate - object). RDF is a framework for describing resources and is suitable for representing semantic web data. For example, "Car1 type Vehicle", "Car1 spatiallyLocatedOn Road1", "Road1 hasTrafficCondition Congested", etc. are all valid RDF triples.

[0164] The specific methods for RDF data generation include:

[0165] (1) Entity representation: Assign a unique identifier to each recognized entity and associate its type using the rdf:type predicate. For example, "Car1 rdf:type Vehicle".

[0166] (2) Relationship representation: Represent the relationships between entities using standard predicates defined in the ontology. For example, "Car1 spatiallyLocatedOn Road1".

[0167] (3) Attribute representation: Represent the attributes of entities using the attribute predicates defined in the ontology and appropriate data types. For example, "Car1 hasColor'red'", "Car1 hasSpeed '60km / h'".

[0168] (4) Confidence Representation: Optionally, add confidence information to each RDF triple to indicate the reliability of the statement. For example, the RDF reification mechanism can be used, such as "Statement1 rdf:subject Car1; rdf:predicate hasColor; rdf:object'red'; hasConfidence 0.95".

[0169] Through the above reasoning steps, the system can infer a large amount of implicit semantic information from the explicit semantic information of the image, greatly enriching the semantic representation of the dataset. These inferred deep semantic information is crucial for evaluating whether the dataset contains the elements required to implement specific software functions.

[0170] Step 500: Extract the necessary conditions for the target function according to the software engineering requirements analysis. Through the requirements analysis method of software engineering, determine the key entities, relationships, and scenario states that need to be identified to implement the target function, and form a list of functional requirements.

[0171] To evaluate whether a dataset is suitable for developing software with specific functions, it is first necessary to clarify the specific requirements and necessary conditions of the function. In the embodiments of this application, the necessary conditions of the target function are systematically extracted through the software engineering requirements analysis method.

[0172] Specifically, software engineering requirements analysis is a systematic process aimed at identifying, documenting, and validating the functional and non-functional requirements of a software system. In the embodiments of this application, the requirements analysis mainly focuses on functional requirements, that is, the specific functions and behaviors that the system needs to implement. Taking an autonomous driving system as an example, the requirements analysis process may include the following steps:

[0173] Step 501: Requirements acquisition

[0174] Collect requirements information related to the target function, which can be obtained through various channels, such as user interviews, expert consultations, research on relevant standards and specifications, etc. For example, for an autonomous driving system, traffic experts can be consulted, traffic regulations can be studied, and the functions of existing autonomous driving systems can be analyzed.

[0175] The specific methods for requirements acquisition include:

[0176] (1) User story analysis: Collect the specific scenarios and behaviors that users expect the system to implement, such as "As a driver, I hope the system can automatically identify the vehicle in front and maintain a safe distance".

[0177] (2) Usage scenario description: Describe in detail the various usage scenarios that the system will face, including normal scenarios and edge scenarios, such as "Driving normally on urban roads", "Navigating under adverse weather conditions", etc.

[0178] (3) Functional objective definition: Clearly define the functional objectives of the system, such as "achieving L3-level autonomous driving", "being able to navigate safely in urban environments", etc.

[0179] (4) Identification of limiting conditions: Identify various limiting conditions of the system, including technical limitations, regulatory limitations, safety requirements, etc., such as "the system must comply with the ISO 26262 functional safety standard".

[0180] Step 502: Requirement analysis

[0181] Analyze the obtained requirement information and decompose it into more specific and actionable functional requirements. This process usually adopts a top-down approach, gradually decomposing high-level requirements into low-level requirements. For example, "safe navigation" can be decomposed into sub-functions such as "obstacle detection", "path planning", "speed control", etc.

[0182] Specific methods for requirement analysis include:

[0183] (1) Functional decomposition: Decompose complex functions into simple and independent sub-functions. For example, decompose "autonomous driving" into three major modules: "perception", "decision-making", and "control", and then further decompose "perception" into "object detection", "lane recognition", "traffic signal recognition", etc.

[0184] (2) Use case analysis: Develop detailed use cases to describe the interaction between the system and the external environment. For example, the use case of "detecting and responding to the braking of the vehicle in front" may include steps such as detecting the vehicle, calculating the relative speed, and determining the braking force.

[0185] (3) Domain model construction: Establish a domain concept model to identify key entities, attributes, and relationships. For example, key entities in the field of autonomous driving include "vehicles", "pedestrians", "roads", "traffic signals", etc., and key relationships include "located at", "approaching", "following", etc.

[0186] (4) State analysis: Identify the possible states and state transitions of the system, especially critical or high-risk states. For example, possible states of an autonomous driving system include "normal driving", "emergency braking", "lane changing", etc.

[0187] Step 503: Extraction of necessary conditions

[0188] Based on the results of requirement analysis, extract the necessary conditions for achieving the target function, including the types of entities, relationship types, and scenario states that the system needs to identify. These necessary conditions directly determine the elements that the dataset needs to contain.

[0189] Specific methods for extracting necessary conditions include:

[0190] (1) Perception requirement analysis: Determine the environmental elements that the system needs to perceive. For example, an autonomous driving system needs to perceive surrounding vehicles, pedestrians, road boundaries, traffic signals, etc.

[0191] (2) Decision requirement analysis: Determine the information required for the system to make decisions. For example, to make a decision on "whether to overtake", the system needs to know information such as the speed of the vehicle in front, whether there is an oncoming vehicle in the oncoming lane, and whether overtaking is allowed on the current road.

[0192] (3) Control requirement analysis: Determine the parameters required for the system to execute control. For example, to achieve smooth lane keeping, the system needs information such as the lane line position, the position and angle of the current vehicle relative to the lane.

[0193] (4) Safety requirement analysis: Determine the key information required to ensure the safe operation of the system. For example, the system needs to be able to detect emergency situations such as suddenly appearing obstacles, sudden braking of the vehicle in front, etc.

[0194] Step 504: Requirement priority division

[0195] Divide the extracted necessary conditions into priorities to distinguish core necessary conditions and secondary conditions. This step helps to focus on the most critical elements when evaluating the dataset. For example, for an autonomous driving system, the detection of vehicles and pedestrians may be core conditions, while identifying the vehicle brand may be a secondary condition.

[0196] Specific methods for requirement priority division include:

[0197] (1) Safety impact assessment: Analyze the degree of impact of each condition on the system safety, and the conditions critical to safety obtain the highest priority. For example, the ability to detect obstacles ahead is crucial for avoiding collisions and should have the highest priority.

[0198] (2) Function impact assessment: Analyze the impact of each condition on the functional integrity of the system, and the conditions for basic functions have a higher priority than those for enhanced functions. For example, the basic path following ability takes precedence over the comfort optimization function.

[0199] (3) Difficulty assessment: Consider the technical difficulty and cost of implementing each condition. Under the same other factors, conditions that are easier to implement can be given priority. For example, navigation on flat roads is more basic and easier to implement than navigation in complex terrains.

[0200] (4) MoSCoW method: Divide the conditions into four categories: Must have, Should have, Could have, and Won't have, to clearly express the priorities.

[0201] Step 505: Form a list of functional requirements

[0202] Organize the extracted necessary conditions into a structured list of functional requirements, clarifying information such as the description, priority, acceptance criteria, etc. for each requirement. This list will serve as the criterion for evaluating the dataset.

[0203] The format of the list of functional requirements may include the following fields, for example:

[0204] (1) Requirement ID: A number that uniquely identifies each requirement, such as "FR-001".

[0205] (2) Requirement category: The functional category to which the requirement belongs, such as "Perception", "Decision-making", "Control", etc.

[0206] (3) Requirement description: Clearly describe the requirement content, such as "The system should be able to detect and track all vehicles on the road".

[0207] (4) Priority: Indicates the importance of the requirement, such as "High", "Medium", "Low" or a numerical identifier (1 - 5).

[0208] (5) Acceptance criteria: Specific and measurable criteria used to verify whether the requirement is met, such as "Under standard lighting conditions, the vehicle detection accuracy rate should be no less than 95%".

[0209] (6) Data requirements: The data characteristics required to implement this requirement, such as "Vehicle images including different angles and different distances".

[0210] Taking an autonomous driving system as an example, the list of functional requirements may include the following:

[0211] ● The system should be able to identify different types of vehicles (sedans, trucks, motorcycles, etc.);

[0212] ● The system should be able to detect pedestrians and predict their movement trajectories;

[0213] ● The system should be able to identify and interpret the status of traffic lights;

[0214] ● The system should be able to detect and track lane lines;

[0215] ● The system should be able to identify road boundaries and obstacles;

[0216] ● The system should be able to judge the driving directions and speeds of other vehicles;

[0217] · The system should be able to detect complex traffic scenarios (such as intersections, roundabouts);

[0218] · The system should be able to work properly under different weather conditions;

[0219] · The system should be able to understand traffic rules and signs.

[0220] Through the above-mentioned requirements analysis process, the necessary conditions for realizing the target function can be systematically extracted, providing clear criteria for subsequent dataset evaluation. These necessary conditions directly reflect the elements that the dataset needs to contain. Only when the dataset fully covers these elements can it support the realization of the target function.

[0221] Step 600: Generate a SPARQL query according to the domain ontology, filter out the images that meet any query condition, and confirm the validity of the target dataset. Generate a SPARQL query according to the functional requirement list and the domain ontology, retrieve the images that meet the requirements, and evaluate the validity of the dataset according to the number and coverage of the images that meet the conditions.

[0222] It should be noted that after completing semantic information extraction, normalization, reasoning, and requirements analysis, the last step is to evaluate whether the dataset contains the elements necessary for realizing the target function. In the embodiment of the present application, the SPARQL query language is used to generate query conditions according to the functional requirement list, filter out the images that meet the conditions, and thus confirm the validity of the dataset.

[0223] Specifically, SPARQL (SPARQL Protocol and RDF Query Language) is a query language for RDF data, capable of searching, adding, modifying, or deleting RDF data. In the embodiment of the present application, the query function of SPARQL is mainly used to find the images that meet specific conditions. The dataset evaluation process may include the following steps, for example:

[0224] Step 601: Query condition construction

[0225] Based on each requirement in the functional requirement list, construct the corresponding SPARQL query condition. These query conditions should accurately reflect the necessary conditions of the requirements and be able to filter out the images that meet the conditions from the RDF data.

[0226] The specific methods for query condition construction include:

[0227] (1) Entity type query: Query the images that contain entities of a specific type. For example, to query the images that contain "vehicle" and "pedestrian", the SPARQL query may be as follows:

[0228]

[0229] (2) Relationship type query: Query the images that contain a specific relationship. For example, to query the images that contain the relationship "vehicle on the road", the SPARQL query may be as follows:

[0230]

[0231]

[0232] (3) Attribute condition query: Query images of entities with specific attributes. For example, to query images containing "red traffic light", the SPARQL query might be as follows:

[0233]

[0234] (4) Scene status query: Query images showing specific scene statuses. For example, to query images showing the "vehicle lane change" scene, the SPARQL query might be as follows:

[0235]

[0236] (5) Composite condition query: A complex query that combines multiple conditions. For example, to query images containing "a vehicle at an intersection and a pedestrian crossing the road", the SPARQL query might be as follows:

[0237] PREFIX rdf:<http: / / www.w3.org / 1999 / 02 / 22-rdf-syntax-ns#>

[0238] PREFIX ad:<http: / / autodrive.org / ontology#>

[0239] SELECT DISTINCT?image

[0240] WHERE{

[0241] ?image rdf:type ad:Image.

[0242] ?intersection ad:appearsIn?image.

[0243] ?intersection rdf:type ad:Intersection.

[0244] ?vehicle ad:appearsIn?image.

[0245] ?vehicle rdf:type ad:Vehicle.

[0246] ?vehicle ad:spatiallyLocatedAt?intersection.

[0247] ?pedestrian ad:appearsIn?image.

[0248] ?pedestrian rdf:type ad:Pedestrian.

[0249] ?crossing ad:appearsIn?image.

[0250] ?crossing rdf:type ad:PedestrianCrossing.

[0251] ?pedestrian ad:performsAction?action.

[0252] ?action rdf:type ad:CrossingAction.

[0253] ?action ad:targetObject?crossing.

[0254] }

[0255] Step 602: Query Execution

[0256] Execute the constructed SPARQL query to retrieve images that meet the conditions from the RDF database. This step requires an efficient query engine, especially when the dataset is large. Common RDF query engines include Apache Jena, RDF4J, etc.

[0257] Specific methods for query execution include:

[0258] (1) Query optimization: Optimize the SPARQL query structure to improve query efficiency. For example, place triples with high selectivity at the front of the query to reduce the scale of intermediate results.

[0259] (2) Batch execution: Organize multiple queries into a batch task to reduce system overhead.

[0260] (3) Parallel execution: Use multi-threading or distributed computing technologies to execute queries in parallel to accelerate the processing of large datasets.

[0261] (4) Incremental execution: For extremely large datasets, an incremental execution approach can be adopted, processing a portion of the data each time to avoid memory overflow.

[0262] Step 603: Result Analysis

[0263] Analyze the query results and evaluate the coverage of each functional requirement by the dataset. This step not only focuses on the existence of images that meet the conditions but also on the quantity, diversity, and quality of these images.

[0264] Specific methods for result analysis include:

[0265] (1) Coverage calculation: Calculate the degree to which each requirement is met. For example, if the requirement is "the system should be able to identify different types of vehicles", then calculate the proportion of images in the dataset that contain images of different types of vehicles such as cars, trucks, motorcycles, etc.

[0266] (2) Diversity analysis: Evaluate whether the images that meet the conditions have sufficient diversity. For example, for the "pedestrian detection" requirement, check whether the dataset contains pedestrian images under different distances, different angles, and different lighting conditions.

[0267] (3) Balance analysis: Check whether the distribution of different category samples is balanced. For example, whether the dataset contains sufficient "minority class" scenarios such as special weather conditions, rare traffic events, etc.

[0268] (4) Edge case analysis: Pay special attention to edge cases or challenging scenarios in functional requirements. For example, the performance under special conditions such as night, bad weather, complex intersections, etc.

[0269] Step 604: Effectiveness evaluation

[0270] Based on the query results and analysis, evaluate the overall effectiveness of the dataset. This evaluation should consider the priority of functional requirements and pay special attention to the satisfaction of core necessary conditions.

[0271] Specific methods for effectiveness evaluation include:

[0272] (1) Weighted scoring: Based on the priority of functional requirements, assign different weights to the satisfaction of each requirement and calculate the weighted total score. For example, core safety requirements may have the highest weight.

[0273] (2) Passing line setting: Set the minimum standard for the effectiveness of the dataset, such as "the coverage rate of all'must-have' requirements is not less than 90%, and there are no completely missing core requirements".

[0274] (3) Graded evaluation: Evaluate the dataset into different grades, such as "excellent", "good", "basically meet", "insufficient", etc., to provide more detailed feedback.

[0275] (4) Gap analysis: Clearly point out the deficiencies of the dataset, such as "lack of pedestrian detection samples under night scenes", to provide specific directions for dataset enhancement.

[0276] Step 605: Report generation

[0277] Generate a detailed dataset effectiveness evaluation report, including overall evaluation results, satisfaction of each functional requirement, advantages and disadvantages of the dataset, etc. This report will guide subsequent dataset enhancement or model training work.

[0278] The specific content of the report generation includes:

[0279] (1) Overall evaluation: A summary evaluation of the overall effectiveness of the dataset, including aspects such as dataset size, coverage, and diversity.

[0280] (2) Requirement satisfaction: List in detail the satisfaction of each functional requirement, including indicators such as coverage rate, number of samples, and sample quality.

[0281] (3) Data distribution analysis: Analyze the distribution of various types of samples in the dataset and identify over-represented or under-represented categories.

[0282] (4) Visualization results: Intuitively display the coverage and deficiencies of the dataset through charts, heatmaps, etc.

[0283] (5) Improvement suggestions: Based on the evaluation results, propose specific dataset enhancement suggestions, such as the types of scenarios and entity types that need to be supplemented.

[0284] Through the above dataset evaluation process, the support ability of the dataset for the target function can be objectively and comprehensively evaluated, providing clear guidance for subsequent dataset enhancement or model training. This evaluation method not only focuses on the surface features of the dataset but also deeply analyzes the semantic content of the dataset to ensure that the dataset can meet the development needs of complex AI systems.

[0285] Through the implementation method provided in this embodiment, first, obtain the target dataset to be evaluated; then, extract the semantic information in the dataset through a pre-trained feature extractor; next, normalize the extracted semantic concepts through a domain ontology and infer the deep semantics; finally, based on the necessary conditions extracted from the software engineering requirements analysis, evaluate whether the dataset contains the elements required to achieve the target function. It can be seen that this application realizes the effectiveness confirmation of the dataset before training, can timely screen out the data that meets the requirements, avoid invalid training, greatly improve the efficiency and quality of AI software development, thereby reducing the development cost and improving the performance of the final product.

[0286] The above has described in detail the preferred specific embodiments of the present invention. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of this application based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. An ontology-based method for validating the effectiveness of a dataset, characterized in that, It includes the following steps: Obtain the target dataset to be evaluated; Extract the semantic information in the target dataset through a pre-trained feature extractor; Normalize the semantic information concepts extracted through the domain ontology and convert them into a standardized semantic representation form; Infer the deep semantics of the semantic information through a semantic reasoner to obtain data in RDF format; Extract the necessary conditions of the target function based on software engineering requirements analysis to provide a clear standard for subsequent dataset evaluation; Generate a SPARQL query according to the domain ontology, filter out the images that meet any query condition, and confirm the validity of the target dataset.

2. The method for validating the effectiveness of a dataset based on ontology according to claim 1, wherein The target dataset includes a large number of images, and each image represents a different scene or situation.

3. The ontology-based dataset validity confirmation method according to claim 1, characterized in that Extract the semantic information in the target dataset through a pre-trained feature extractor. Specifically, use the Faster R-CNN model as the pre-trained feature extractor to extract the semantic information of each image in the target dataset in the way of scene graph generation / detection, and the output results include entities and entity-relationship-entity triples.

4. The method for validating the effectiveness of a dataset based on ontology according to claim 3, wherein, The processing flow of the Faster R-CNN model includes the following steps: Apply a convolutional neural network to the input images in the target dataset to extract feature maps, including scene graphs; Generate candidate regions through the Region Proposal Network, and filter out candidate regions with low overlap and high scores through the non-maximum suppression algorithm; Apply RoI pooling to each candidate region to map candidate regions of different sizes into a fixed-size feature representation; Perform object classification and bounding box regression on the pooled features to determine the object category and precise location; Analyze the spatial and semantic relationships between objects to generate entity-relationship-entity triples; Integrate the object and the relationship information between them to form a complete scene graph.

5. The method for validating the effectiveness of a dataset based on an ontology according to claim 1, characterized in that, Normalize the semantic information concepts extracted through the domain ontology and convert them into a standardized semantic representation form. Specifically, use the domain ontology to normalize the semantic information output by the feature extractor into a standardized semantic representation form.

6. The method for validating the effectiveness of a dataset based on ontology according to claim 5, characterized in that The normalization of semantic information concepts includes the following steps: Map the entities recognized by the feature extractor to the standard concepts defined in the domain ontology; Map the relationships recognized by the feature extractor to the standard relationships defined in the domain ontology; Standardize the various attributes of the entities recognized by the feature extractor and map them to the standard attributes defined in the domain ontology; Process the uncertain information output by the feature extractor, and finally convert the original semantic information into a standardized semantic representation that conforms to the ontology.

7. The method for validating the effectiveness of a dataset based on ontology according to claim 1, wherein, Infer the deep semantics of the semantic information through a semantic reasoner to obtain data in RDF format. Specifically, use the HermiT reasoner to infer deep semantic relationships based on the normalized semantic concepts and the inference rules defined in the domain ontology to generate semantic data in RDF format.

8. The method for validating the effectiveness of a dataset based on ontology according to claim 7, characterized in that, The RDF format is a triple format, including subject-predicate-object; the specific method for generating the semantic data in RDF format includes entity representation, relationship representation, attribute representation, and confidence representation.

9. The method for validating the effectiveness of a dataset based on an ontology according to claim 1, wherein Extract the necessary conditions for the target function based on the requirements analysis of software engineering, providing clear criteria for subsequent dataset evaluation. By using the requirements analysis method of software engineering, identify the key entities, relationships, and scenario states required to achieve the target function, and form a list of functional requirements, specifically including: Collect requirement information related to the target function, analyze the obtained requirement information, and decompose it into more specific and operable functional requirements; Based on the results of requirements analysis, extract the necessary conditions for achieving the target function, including the entity types, relationship types, and scenario states that the system needs to identify; Prioritize the extracted necessary conditions to distinguish core necessary conditions from secondary conditions; Organize the extracted necessary conditions into a structured list of functional requirements, clarifying the description, priority, and acceptance criteria information for each requirement; this list will serve as the criteria for evaluating the dataset.

10. The method for validating the effectiveness of a dataset based on ontology according to claim 1, characterized in that Generate SPARQL queries according to the domain ontology, filter out the images that meet any query condition, and confirm the validity of the target dataset. Specifically, generate SPARQL queries based on the list of functional requirements and the domain ontology, retrieve the images that meet the requirements, and evaluate the validity of the dataset according to the number and coverage of the images that meet the conditions.