Short message AI training model intelligent matching method and system

By acquiring and analyzing natural language descriptions and file features, combining the model feature image library and dynamic weight matching algorithm, the problem of insufficient comprehensive and accurate model matching in the existing technology is solved, and efficient and accurate intelligent matching of SMS AI models is achieved, and self-optimization capabilities are achieved.

CN119939199AActive Publication Date: 2025-05-06BEIJING YULORE INNOVATION TECH

Patent Information

Application Number
CN202510436872.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the intelligent matching of SMS AI models, the existing technology has problems such as insufficient input of a single natural language, lack of standardized model capability description system, and static matching strategies, resulting in insufficient matching results, low accuracy and inability to self-optimize.

Method used

By obtaining natural language descriptions, file names, file contents and file formats, semantic chunking and feature extraction are performed to form structured instructions and multi-dimensional feature data. The preset model feature image library and feature quadruple dynamic weight matching algorithm are used for processing, the candidate AI training model is output, and the weight parameters and model feature image library are updated through user feedback.

Benefits of technology

More comprehensive and accurate model matching is achieved, and through multi-dimensional feature analysis and dynamic weight adjustment, the input quality and accuracy of matching are improved, manual intervention is reduced, and the ability to evolve is possessed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939199A_ABST
    Figure CN119939199A_ABST
Patent Text Reader

Abstract

The invention provides a short message AI training model intelligent matching method and system. The method comprises the steps of forming a structured instruction and multi-dimensional feature data; according to a preset model feature portrait library, processing the structured instruction and the multi-dimensional feature data by adopting a feature tetrad dynamic weight matching algorithm, and outputting a candidate AI training model; according to the candidate AI training model, constructing an intelligent auxiliary decision-making interface, analyzing and displaying a processing logic flow chart, a feature analysis report and a historical use record of the candidate AI training model, and generating a user feedback result; and using a user feedback result to update a weight parameter in the feature tetrad dynamic weight matching algorithm, performing re-evaluation and classification on the AI training model in the preset model feature portrait library according to the updated weight parameter, and generating an optimized model feature portrait library. According to the invention, the data processing efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a method and system for intelligent matching of SMS AI training models. Background Art

[0002] The field of SMS AI model intelligent matching involves multiple technical directions such as natural language processing, feature recognition, model matching, etc. This field aims to achieve accurate matching of data processing requirements with AI model functions and improve data processing efficiency and accuracy.

[0003] At present, common technical solutions mainly use fixed rule matching or simple keyword matching. For example, manual selection is performed based on a preset model classification directory, or literal matching is performed between keywords in the task description and model labels. Although this method is simple to implement, it cannot meet the precise matching requirements in complex scenarios. It often requires professionals to make multiple attempts and adjustments, which is inefficient.

[0004] The most similar existing technology uses a model recommendation solution based on semantic understanding, which performs semantic analysis on the natural language description input by the user, calculates similarity with the model function description library, and recommends potentially applicable AI models. Its technical principle is to construct a model function vector space, map user requirements to the space for nearest neighbor search, and thus achieve model matching.

[0005] However, this technology still has the following major problems: First, it relies solely on a single natural language input for matching and fails to fully utilize multi-dimensional information such as file features, resulting in incomplete matching results; second, it lacks a standardized model capability description system and cannot accurately quantify model capabilities, resulting in low matching accuracy; third, it adopts a static matching strategy and cannot be dynamically optimized and adjusted according to actual usage results, and the system cannot evolve by itself. Summary of the invention

[0006] In view of this, the present application provides a method and system for intelligent matching of SMS AI training models, which solves the problems in the prior art of relying only on a single natural language input, lacking a standardized model capability description system, and adopting a static matching strategy.

[0007] The present application embodiment provides a method for intelligent matching of SMS AI training models, including: Obtaining a natural language description, a file name, a file content, and a file format of the data to be processed, and performing semantic segmentation using the natural language description to form structured instructions, and performing feature extraction on the file name, the file content, and the file format to form multi-dimensional feature data; According to the preset model feature portrait library, a feature quadruple dynamic weight matching algorithm is used to process the structured instructions and the multi-dimensional feature data, and a candidate AI training model is output; Based on the candidate AI training model, an intelligent auxiliary decision-making interface is constructed to analyze and display the processing logic flow chart, feature analysis report and historical usage record of the candidate AI training model to generate user feedback results; The user feedback results are used to update the weight parameters in the feature quadruple dynamic weight matching algorithm, and the AI ​​training model in the preset model feature portrait library is re-evaluated and classified according to the updated weight parameters to generate an optimized model feature portrait library.

[0008] Optionally, the use of the natural language description to perform semantic segmentation to form structured instructions includes: dividing the natural language description into multiple semantic blocks by calling a pre-trained scenario-based intent recognition model; identifying and extracting the core action instructions, target objects and data carriers of each semantic block to obtain preliminary semantic information containing the core action instructions, target objects and data carriers; based on the preliminary semantic information, mapping the core action instructions, the target objects and the implementation data carriers to a preset intent knowledge graph to form the structured instructions.

[0009] Optionally, feature extraction is performed on the file name, the file content and the file format to form multi-dimensional feature data, including: calling a dynamic rule engine to parse the file name, extract time, content, status and format information, and obtain file name features; performing first row detection on the file content to identify column header features, and using random sampling to determine the data pattern, performing format recognition on specific fields, and generating file content features; parsing the file format according to a preset format decoding matrix, determining the processing method for encrypted files, compressed packages or unstructured data, and forming file format features; and using a feature fusion algorithm to process the file name features, file content features and file format features to form the multi-dimensional feature data.

[0010] Optionally, the preset model feature portrait library is constructed through standardized description, including: obtaining the functional data of the initial candidate AI training model, and using the scenario tree structure to describe the applicable scenario, and determining the specific application field and processing target of the initial candidate AI training model; according to the technical specifications of the initial candidate AI training model, extracting technical parameters including required fields, data format and data volume restrictions to form input requirement characteristics; using standardized capability description language to process the functional characteristics of the initial candidate AI training model, breaking down specific functions into measurable processing units, and generating processing capability characteristics; integrating the specific application field and the processing target, the input requirement characteristics and the processing capability characteristics to generate a three-dimensional feature file of the candidate AI training model as a preset model feature portrait library.

[0011] Optionally, the execution process of the feature quadruple dynamic weight matching algorithm includes: using the first deep learning model to calculate the semantic similarity between the structured instruction and the functional description of the candidate AI training model, and generating a natural language weight score; processing the key information of the file name feature through regular expressions and rule engines, and matching it with the target data type of the candidate AI training model to form a file name feature weight score; calculating the matching degree between the file content features and the input requirements of the candidate AI training model, and outputting the content feature weight score; analyzing the data format conversion cost corresponding to the file format features to generate a format adaptation weight score; using a preset weight coefficient to perform weighted calculation on the natural language weight score, file name feature weight score, content feature weight score and format adaptation weight score to obtain a comprehensive matching degree, and using the comprehensive matching degree as the final scoring result of the candidate AI training model.

[0012] Optionally, before executing the feature quadruple dynamic weight matching algorithm, it also includes: collecting historical matching records, identifying high-frequency typical subtask combinations through data mining, building a high-frequency subtask optimization pool, and outputting task combination patterns; using the task combination pattern to analyze the execution process of each high-frequency subtask, extracting data dependencies, execution order requirements and resource consumption constraints, and generating a set of constraint equations; processing the search space of candidate AI training models according to the set of constraint equations, performing pruning operations, and forming an optimized set of candidate AI models.

[0013] Optionally, the task combination pattern is used to analyze the execution process of each high-frequency subtask, extract data dependencies, execution order requirements and resource consumption constraints, and generate a set of constraint equations, including: calling the first deep learning model to process the model combination pattern in historical successful cases, and outputting a model link prediction strategy through pattern recognition and feature learning; analyzing the high-frequency subtasks according to the model link prediction strategy, extracting data flow rules and computing resource limitation constraint elements, and forming formalized constraint conditions; converting the constraint conditions into a computational model, and establishing a quantifiable set of constraint equations to guide subsequent model search and optimization processes.

[0014] Optionally, the intelligent matching method also includes: adopting a feature compression mechanism in the feature quadruple dynamic weight matching algorithm, including: obtaining original feature space data, constructing a feature pyramid structure through multi-level feature extraction, and outputting feature combinations of different granularity levels; based on the feature combination, performing top-down iterative matching operations, using compressed feature data to screen candidate AI training models, and generating preliminary matching results; based on the preliminary matching results, constructing a feature weight adjustment algorithm, dynamically optimizing the feature matching process, and outputting an optimized weight matching algorithm.

[0015] Optionally, the top-down iterative matching operation is used to screen candidate AI training models using compressed feature data, including: performing preliminary screening through the highest-level compressed features of the feature pyramid structure, determining the category range of candidate AI training models, and outputting model category data; using the model category data, performing feature matching refinement at each level of the feature pyramid in turn, determining the candidate AI training models through layer-by-layer screening, and generating screening results; evaluating and verifying the screening results, storing the feature weight combinations and matching strategies whose screening effects reach a preset threshold in the algorithm library, and forming a reusable feature matching solution.

[0016] Correspondingly, the present invention also provides an intelligent matching system for SMS AI training models, including: a feature acquisition module, which is used to obtain the natural language description, file name, file content and file format of the data to be processed, and use the natural language description to perform semantic segmentation to form structured instructions, and extract features from the file name, the file content and the file format to form multi-dimensional feature data; a model matching module, which is used to process the structured instructions and the multi-dimensional feature data according to a preset model feature portrait library using a feature quadruple dynamic weight matching algorithm, and output a candidate AI training model; a decision interface module, which is used to construct an intelligent auxiliary decision interface according to the candidate AI training model, and generate user feedback results by analyzing and displaying the processing logic flow chart, feature analysis report and historical usage record of the candidate AI training model; an optimization and update module, which is used to update the weight parameters in the feature quadruple dynamic weight matching algorithm using the user feedback results, re-evaluate and classify the AI ​​training models in the preset model feature portrait library according to the updated weight parameters, and output the optimized model feature portrait library.

[0017] An embodiment of the present application also provides a computer system, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned SMS AI training model intelligent matching method.

[0018] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the above-mentioned SMS AI training model intelligent matching method.

[0019] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-mentioned SMS AI training model intelligent matching method.

[0020] This application has the following technical effects: by constructing a multi-dimensional input parsing system, data feature information is fully captured, and the input quality of model matching is improved; by utilizing the feature quadruple dynamic weight matching algorithm, the accuracy of model matching is self-optimized, reducing human intervention; a standardized three-dimensional archive system for model feature portraits is designed, which effectively quantifies the model capability boundary and improves the matching accuracy; a feedback learning closed-loop mechanism is established, and the matching effect is continuously optimized through human-computer collaboration, realizing the system's self-evolution capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can also be obtained based on these drawings without creative work.

[0022] Figure 1 A flowchart of the SMS AI training model intelligent matching method provided in the embodiment of the present application; Figure 2 A schematic diagram of a multi-dimensional feature analysis process provided in an embodiment of the present application; Figure 3 A schematic diagram of the calculation flow of the feature quadruple dynamic weight matching algorithm provided in the embodiment of the present application; Figure 4 A schematic diagram of an intelligent decision-making support interface provided in an embodiment of the present application; Figure 5 This is a structural block diagram of the SMS AI training model intelligent matching system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the technical scheme in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present disclosure.

[0024] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0025] The term "and / or" herein only describes an association relationship, indicating that three relationships may exist. For example, A and / or B may represent the following three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set consisting of A, B, and C.

[0026] like Figure 1 As shown, the embodiment of the present application provides a method for intelligent matching of SMS AI training models, including: S1: Obtain a natural language description, file name, file content and file format of the data to be processed, and use the natural language description to perform semantic segmentation to form structured instructions, and form multi-dimensional feature data by extracting features of the file name, the file content and the file format.

[0027] This step builds a comprehensive input feature parsing framework. In actual implementation, the embodiment of the present application first receives the natural language description and the data file to be processed provided by the user. For the processing of natural language descriptions, the embodiment of the present application calls a pre-trained scenario-based intent recognition model to divide complex user descriptions into semantic blocks containing core semantics.

[0028] For example, when a user inputs "I need to process a batch of SMS numbers and check their validity", the embodiment of the present application will decompose it into semantic units such as "I need", "process a batch of SMS numbers" and "check their validity".

[0029] Subsequently, the embodiment of the present application identifies and extracts the core action instructions (such as "processing", "checking"), target objects (such as "SMS number"), and data carriers (implied as files) from these semantic blocks to form preliminary semantic information. The embodiment of the present application maps these semantic elements to the pre-built intent knowledge graph, converting natural language into structured instructions that can be processed by computers, such as "Action: ValidateObject: SMS_NumberProperty: Validity".

[0030] At the same time, the embodiment of the present application performs multi-dimensional feature extraction on the uploaded file. The embodiment of the present application calls the dynamic rule engine to parse the file name, extracting time information (such as "2024Q3" for the third quarter of 2024), content information (such as "number list" for data type), status information (such as "FINAL version" for the final version) and format information.

[0031] This information provides important clues for understanding the business background of the data. The embodiment of the present application uses progressive intelligent sniffing technology for the file content, identifies the column header features through the first line detection, and understands the structural organization of the data; then analyzes the distribution characteristics of the data through random sampling; and finally performs special format recognition for specific fields (such as SMS numbers, content, etc.).

[0032] For example, the embodiment of the present application can automatically identify the format type, distribution characteristics and possible abnormal patterns of the SMS number. In addition, the embodiment of the present application also parses the file format according to the preset format decoding matrix to determine whether special processing such as decryption, decompression or OCR is required.

[0033] The embodiment of the present application integrates these multi-dimensional features into a unified feature representation through a feature fusion algorithm, providing rich input information for subsequent model matching. This multi-dimensional feature acquisition method solves the limitation of the prior art that only relies on a single dimension for matching, enabling the embodiment of the present application to more comprehensively understand user needs and data features.

[0034] like Figure 2 As shown, the natural language description is used to perform semantic segmentation to form structured instructions, including: A1: The natural language description is divided into multiple semantic blocks by calling a pre-trained scenario-based intent recognition model.

[0035] First, a natural language description input by the user is received, such as "I need to analyze this batch of SMS number data and check its validity, and extract invalid numbers to generate a report." For such complex descriptions, the embodiment of the present application calls a pre-trained scenario-based intent recognition model for processing.

[0036] The model is based on pre-trained language models such as improved BERT or RoBERTa, and uses attention mechanism and semantic understanding ability to identify semantic boundaries and turning points in sentences.

[0037] The embodiment of the present application will split the original description into multiple independent semantic blocks according to the semantic integrity, for example, the above description is split into three semantic blocks of "analyzing SMS number data", "checking number validity" and "extracting invalid numbers to generate reports". This semantic block processing enables the embodiment of the present application to decompose complex, multi-task descriptions into basic units that can be understood and processed independently, laying the foundation for subsequent accurate intent recognition.

[0038] In practical applications, the embodiments of the present application will also consider domain-specific language habits and expressions, such as professional terms and common expression patterns in the field of SMS data processing, and improve the accuracy of segmentation through domain adaptability fine-tuning. In addition, the embodiments of the present application can also handle descriptions of implicit relationships, such as the user may not explicitly indicate the source of the data, but the embodiments of the present application can understand it implicitly in the uploaded file through context. This advanced semantic segmentation capability greatly improves the ability of the embodiments of the present application to understand complex natural language instructions, laying the foundation for subsequent accurate matching.

[0039] A2: Identify and extract the core action instructions, target objects and data carriers of each semantic block, and obtain preliminary semantic information including the core action instructions, target objects and data carriers.

[0040] The embodiment of the present application adopts structured semantic analysis technology to perform in-depth analysis on each semantic block.

[0041] Specifically, the embodiment of the present application uses semantic role labeling technology to identify the core elements in each semantic block: action instructions (such as "analysis", "check", "extract", "generate") represent the operations that the user wants to perform; target objects (such as "SMS number data", "validity", "invalid number", "report") represent the recipients of the operation; data carriers (such as implicit "upload files" or explicitly mentioned "Excel tables") represent the storage form of data. In order to improve the accuracy of extraction, the embodiment of the present application constructs a domain-specific action dictionary and object ontology library, which contains common operation types and data objects in the field of SMS data processing.

[0042] For example, for the semantic block "check number validity", the embodiment of the present application identifies "check" as an action instruction (classified as a "verification" type operation), "number validity" as a target object (including the subject "number" and the attribute "validity"), and the data carrier is inferred from the context as an uploaded file. The embodiment of the present application can also handle compound actions and nested objects, such as "extract invalid numbers and generate reports" contains two consecutive actions of "extract" and "generate", and two related objects of "invalid number" and "report". Through this structured semantic parsing, the embodiment of the present application converts natural language descriptions into preliminary semantic information with a clear subject-predicate-object structure, preparing for subsequent standardization processing.

[0043] A3: Based on the preliminary semantic information, the core action instructions, the target object and the implementation data carrier are mapped to a preset intention knowledge graph to form the structured instructions.

[0044] The embodiment of the present application realizes the conversion from preliminary semantic information to standardized structured instructions. The embodiment of the present application constructs a complete intention knowledge graph, which contains standardized representations of common operations in the field, hierarchical classifications of object types, and legal relationships between them.

[0045] For example, the graph defines that "validation" operations can be applied to the "validity", "format", "status" and other properties of "number" objects. After the embodiment of the present application obtains preliminary semantic information, the action instructions, target objects and data carriers therein are mapped to the standard nodes of the knowledge graph.

[0046] This process involves semantic equivalence judgment, such as mapping different expressions such as "check", "verify", and "confirm" to the same "Validate" standard operation.

[0047] At the same time, the embodiments of the present application will also handle the relationship between synonyms, hypernyms and hyponyms, such as mapping "mobile phone number", "telephone number", etc. to the "TelephoneNumber" standard object type.

[0048] Through this mapping, the embodiment of the present application converts various forms of natural language descriptions into a unified structured instruction format: Such as "Action: ValidateObject: TelephoneNumberProperty: Validity".

[0049] This standardized structured instruction eliminates the ambiguity and diversity of natural language expressions, enabling the embodiments of the present application to accurately understand user intent and accurately match it with the functional description in the model library.

[0050] In addition, when the mapping process encounters a new concept that is not defined in the knowledge graph, the embodiment of the present application will also start the concept expansion mechanism, find the closest known concept through semantic similarity calculation, or create a new concept node to keep the knowledge graph updated and expanded. This semantic standardization mechanism based on the knowledge graph is the key link for the embodiment of the present application to accurately understand the user's intention, and provides a reliable foundation for subsequent model matching.

[0051] Among them, the processing of natural language description is based on scenario-based intent recognition technology. The embodiment of the present application first divides the user's description text into multiple semantic blocks, such as "I need + process a batch of SMS number data + check its validity". This block processing helps to understand the user's expression intention more accurately. Then, core elements are extracted from these semantic blocks, including action instructions ("process", "check"), target objects ("SMS number data") and data carriers (implied as files). This structured analysis method can standardize the same intentions in different expressions and improve the accuracy of subsequent matching.

[0052] The embodiment of the present application maps the extracted semantic elements to a pre-built intent knowledge graph. The knowledge graph contains standardized representations of common operations in the field, and through mapping operations, natural language is converted into structured instructions that can be processed by computers. For example, "check the validity of the SMS number" may be mapped to a structured representation such as "Action: ValidateObject: SMS_NumberProperty: Validity".

[0053] In addition, by extracting features from the file name, the file content, and the file format, multi-dimensional feature data is formed, including: B1: calling the dynamic rule engine to parse the file name, extract time, content, status and format information, and obtain file name features.

[0054] For the extraction of file features, the embodiment of the present application adopts a multi-level feature analysis strategy: For file names, the embodiment of the present application uses a dynamic rule engine for parsing. The engine contains parsing rules for different naming patterns, and can extract multi-dimensional information from the file name, such as time features (such as "2024Q3" for the third quarter of 2024), content features (such as "number details table" for data types), and status features (such as "FINAL version" for file status).

[0055] The embodiment of the present application implements an intelligent file name parsing mechanism. File names usually contain rich implicit information, which is of great value for understanding the business background of data. The embodiment of the present application constructs a dynamic rule engine, which includes multiple sets of parsing rules for different naming conventions, and can adaptively identify and extract key information in file names.

[0056] The embodiment of the present application will identify the time information in the file name, such as "2024Q3" for the third quarter of 2024, "20240315" for March 15, 2024, and "Jan-Feb" for January to February. The embodiment of the present application extracts the content information in the file name, such as "SMS number list", "Customer information table", "certain record", etc., to identify the basic type and purpose of the data.

[0057] At the same time, the embodiment of the present application will also analyze status information, such as "FINAL", "V2", "Draft", etc., to understand the version and completion level of the file. The embodiment of the present application will identify format suffixes and special identifiers to determine the technical format of the file and possible special processing requirements. In order to improve the accuracy of parsing, the dynamic rule engine adopts a hierarchical rule organization structure, from general rules to specific domain rules, trying to match layer by layer, and can learn and remember new naming patterns.

[0058] For example, when processing a file named "2024Q3_SMS_NumList_Corp_FINAL.xlsx", the embodiment of the present application can extract the time "2024 Q3", the content "corporate SMS number list", the status "final version", and the format "Excel table". These features extracted from the file name provide important clues for understanding the business background and usage scenarios of the data, and play a key role in subsequent model matching.

[0059] B2: Perform first-row detection on the file content to identify column header features, use random sampling to determine data patterns, perform format recognition on specific fields, and generate file content features.

[0060] For the file content, the embodiment of the present application adopts progressive intelligent sniffing technology. First, the column header features are identified through the first line detection to understand the structural organization of the data; then the distribution characteristics of the data are analyzed through random sampling; finally, special format recognition is performed for specific fields (such as SMS numbers, content, etc.). This multi-level content analysis ensures a comprehensive grasp of the file content characteristics.

[0061] The embodiment of the present application adopts a multi-level intelligent content analysis strategy. The embodiment of the present application performs first line detection, analyzes the header or column name information of the file, and understands the basic structure and organization of the data.

[0062] The embodiment of the present application will identify key field names, such as "mobile phone number", "text message content", "sending status", etc., and preliminarily determine the field and purpose of the data. In order to cope with different header naming habits, the embodiment of the present application constructs a field-specific field name synonym library that can recognize the same semantics under different expressions. The embodiment of the present application analyzes the data content through random sampling technology to understand the data type, value range and distribution characteristics of each field.

[0063] For example, the embodiment of the present application can identify that a column contains a mobile phone number in an 11-digit format, another column contains text content within 200 words, and another column contains date and time information. Through this distributed sampling analysis, the embodiment of the present application can quickly understand the overall characteristics of a large data file without having to fully scan the entire content.

[0064] Subsequently, the embodiment of the present application performs in-depth format recognition for specific fields, especially special analysis of key fields. For example, for the SMS number field, the embodiment of the present application will detect whether its format complies with domestic mobile phone number specifications, whether it contains international area codes, whether it is mixed with landline numbers, etc.; for SMS content, the embodiment of the present application will analyze text length distribution, keyword frequency, language type and other characteristics.

[0065] The embodiment of the present application combines these analysis results to generate a structured file content feature description, including data structure information (such as the number and type of fields), content feature information (such as data distribution, abnormal patterns) and domain feature information (such as whether it is marketing SMS data, whether it is a notification SMS, etc.). This comprehensive and in-depth content analysis provides an accurate data feature description for subsequent model matching, which is an important basis for achieving accurate matching.

[0066] B3: According to a preset format decoding matrix, the file format is parsed, and a processing method for encrypted files, compressed packages or unstructured data is determined to form file format features.

[0067] For file formats, the present application embodiment constructs a format decoding matrix that can identify and process various common formats. For encrypted files, compressed packages or unstructured data (such as PDF, pictures), the present application embodiment will start the corresponding processing flow to ensure that valid information can be extracted.

[0068] The embodiment of the present application realizes an intelligent file format processing mechanism. The embodiment of the present application constructs a decoding matrix containing various common data formats, which can not only judge the format by the file extension, but also verify the format by the file header information and content characteristics, solving the problem that the extension may be modified.

[0069] Based on the identified file format, the embodiments of the present application adopt differentiated processing strategies: for standard structured data formats (such as CSV, Excel, SQL database files, etc.), the embodiments of the present application directly call the corresponding parser to extract the data structure and content; for encrypted files, the embodiments of the present application will detect the encryption type, and if the user provides access credentials, it will automatically call the decryption module for processing; for compressed packages (such as ZIP, RAR, etc.), the embodiments of the present application will analyze the content structure of the compressed package, automatically decompress and recursively process the internal files; for unstructured data such as PDF or pictures, the embodiments of the present application will start the OCR text extraction process to convert the tables and text in the image into processable structured data.

[0070] In addition, the embodiments of the present application can also process composite format files, such as Excel files containing multiple worksheets or PDF documents containing multiple media types. Through this intelligent format processing, the embodiments of the present application form a detailed description of file format features, including file type, structural complexity, parsing difficulty, and possible preprocessing operations. These format features are crucial for the subsequent selection of appropriate processing models, especially when evaluating data conversion costs and processing compatibility, format features can provide key decision-making basis.

[0071] B4: Using a feature fusion algorithm to process the file name features, file content features and file format features to form the multi-dimensional feature data.

[0072] The embodiment of the present application integrates feature information of different dimensions into unified multi-dimensional feature data through a feature fusion algorithm. The fusion process is not a simple feature splicing, but takes into account the relevance and importance between different features to form a more expressive feature representation.

[0073] Specifically, when processing a file named "2024Q3_Customer SMS Number List_V2.xlsx", the embodiment of the present application will parse it from multiple dimensions: first, it will identify that this is an Excel file (format feature), containing customer SMS number data (content feature) in the third quarter of 2024 (time feature), and it is the second version (status feature); then, by analyzing the file content, it is found that it contains fields such as mobile phone number, operator, status, etc. (content structure feature); finally, combined with the user's natural language description "need to check the validity of this batch of numbers", a complete task understanding is formed. This multi-dimensional feature acquisition lays a solid foundation for subsequent precise model matching.

[0074] The embodiment of the present application implements an advanced feature integration mechanism. Feature fusion is not a simple feature splicing, but takes into account the correlation and complementarity between different features to form a more expressive unified feature representation. The embodiment of the present application uses feature association analysis to identify the mutual support or conflict relationship between features of different dimensions.

[0075] For example, when the file name contains "SMS number list" and content analysis shows that there is indeed a mobile phone number field, the two features support each other and increase the confidence of the relevant features; when the file name implies that it is the "final version" but content analysis finds a large number of missing values, this may indicate that the file status is inconsistent with the content, and the embodiment of the present application will adjust the feature weights accordingly.

[0076] The embodiment of the present application implements a feature supplementation mechanism, using high-confidence features in one dimension to infer missing features in another dimension. For example, when the file content is clearly displayed as SMS data, but the file name does not contain relevant information, the embodiment of the present application will automatically supplement the content type feature.

[0077] At the same time, the embodiment of the present application also performs feature unification processing, converting features of different dimensions into a consistent representation format to facilitate subsequent matching calculations. In addition, the embodiment of the present application implements multi-scale feature extraction, which not only retains fine-grained specific features (such as specific field names and formats), but also generates coarse-grained summary features (such as the overall type and field of data). Through this multi-level feature fusion processing, the embodiment of the present application ultimately generates structured multi-dimensional feature data, which comprehensively describes the various aspects of the data to be processed, and provides rich input information for subsequent precise model matching. This feature fusion method solves the problem of one-sided understanding caused by relying on a single feature dimension in traditional methods, and significantly improves the depth and comprehensiveness of the understanding of data features in the embodiment of the present application.

[0078] S2: Based on the preset model feature portrait library, a feature quadruple dynamic weight matching algorithm is used to process the structured instructions and the multi-dimensional feature data, and a candidate AI training model is output.

[0079] This step implements the core model matching function. The embodiment of the present application constructs a standardized model feature portrait library. The library uses a three-dimensional modeling method to describe the capability boundary of each AI training model. The applicable scenario dimension is described using a multi-level classified scenario tree structure, such as "data cleaning> phone number verification> number validity detection", to accurately locate the application field of the model. The input requirement dimension defines detailed data specifications, including technical parameters such as required fields, data formats, and data volume limits, and clarifies the prerequisites for the model to process data. The processing capability dimension uses a standardized capability description language to refine the model's functions into measurable processing units, such as accuracy, processing speed, supported abnormal situations, etc. The features of these three dimensions are integrated into a complete feature file of the model through a unified representation method, forming a model feature portrait library. This standardized model description system solves the problem that model capabilities are difficult to accurately describe in the prior art, and provides a basis for accurate matching.

[0080] On this basis, the embodiment of the present application implements a feature quadruple dynamic weight matching algorithm. The algorithm decomposes the matching process into four dimensions: natural language matching, file name feature matching, content feature matching, and format adaptation matching. The embodiment of the present application uses a deep learning model to calculate the semantic similarity between structured instructions and model function descriptions to generate a natural language weight score. This semantic-based matching method can better understand the user's true intentions than simple keyword matching.

[0081] The embodiment of the present application processes the file name features through the rule engine, analyzes the matching degree between the file name features and the target data type of the candidate model, and forms a file name feature weight score. Then, the embodiment of the present application calculates the matching degree between the file content features and the model input requirements, considers factors such as data structure and field type, and outputs the content feature weight score.

[0082] At the same time, the embodiment of the present application also analyzes the compatibility of the file format with the model-supported format, evaluates the possible conversion cost, and generates a format adaptation weight score. The embodiment of the present application uses a preset weight coefficient (such as 40% for natural language, 20% for file name, 30% for content, and 10% for format) to perform a weighted calculation on these four scores and obtain a comprehensive matching degree as the final scoring result of the candidate model.

[0083] In the embodiment of the present application, the first deep learning model adopts a hierarchical hybrid architecture design, which is specifically used to calculate the semantic similarity between structured instructions and model function descriptions. The base layer of the model adopts a pre-trained BERT-base model (12-layer Transformer encoder, hidden layer dimension 768, 12 attention heads), and is further pre-trained on a corpus in the field of SMS processing.

[0084] The domain adaptation pre-training used about 2 million text data related to SMS processing, including task descriptions, processing instructions, and result reports. The pre-training adopted a dual-task method of masked language model and next sentence prediction, with a learning rate of 2e-5, a batch size of 32, and 8 epochs. On this basis, the model added an attention enhancement layer specifically for structured instruction processing. This layer uses a multi-head cross-attention mechanism (8 attention heads) to enable the model to better capture the correspondence between structured instructions and natural language descriptions. The top layer of the model is a specific task fine-tuning layer, which contains two fully connected layers (with dimensions of 768 and 256, respectively), uses the GELU activation function, and finally calculates the output semantic matching score through cosine similarity.

[0085] The model training adopts a two-stage strategy: first, supervised learning is performed on a labeled semantic similarity dataset (about 100,000 pairs of labeled samples), using the AdamW optimizer, a learning rate of 3e-5, and a weight decay of 0.01; then online fine-tuning is performed through a feedback learning mechanism, using user feedback as a signal and a contrastive learning method to further optimize the model performance. Evaluation indicators include semantic matching accuracy, ranking quality (NDCG@5), and consistency score. On the internal test set, the model achieved a matching accuracy of 89.7%, an increase of 15.3 percentage points over the baseline model. In addition, the model also implements incremental learning capabilities, which can adapt to new matching patterns and domain terms without complete retraining, maintaining long-term effectiveness.

[0086] In order to improve the matching efficiency, the embodiment of the present application also implements two important optimization mechanisms. On the one hand, the embodiment of the present application identifies the typical subtask combinations that appear frequently through data mining, such as the task chain of "mobile phone number verification → address normalization → customer classification", and extracts key constraints for these high-frequency subtasks to form a set of constraint equations for pre-pruning the search space to improve matching efficiency.

[0087] On the other hand, the embodiment of the present application introduces a feature compression mechanism, constructs a feature pyramid structure, and implements a top-down iterative matching strategy. This strategy first uses highly compressed features for rough screening to quickly determine the possible model categories; then uses more detailed features layer by layer for refined screening, and finally obtains the most suitable model. These two optimization mechanisms greatly improve the matching speed and accuracy of the embodiment of the present application, especially when dealing with complex tasks and a large number of candidate models.

[0088] like Figure 3 As shown, the preset model feature portrait library is constructed through standardized description, including: C1: Obtain the functional data of the initial candidate AI training model, and use the scenario tree structure to describe the applicable scenario, and determine the specific application field and processing target of the initial candidate AI training model.

[0089] The embodiments of the present application implement a standardized model capability description mechanism. The embodiments of the present application collect the original functional description of the AI ​​training model, including development documents, technical specifications, instructions for use and other materials. These original descriptions usually exist in the form of unstructured text, with inconsistent content, and are difficult to use directly for precise matching.

[0090] To address this problem, the embodiment of the present application constructs a multi-level scenario tree structure as a framework for the standardized description model applicable scenarios. The scenario tree is organized in a hierarchical classification manner, with each layer being refined from the top-level categories to the bottom-level specific application scenarios. For example, the top layer may be "SMS data processing", the second layer is subdivided into "data cleaning", "content analysis", "number management", etc., and the third layer is further refined into specific scenarios such as "number validity detection", "number location identification", and "operator classification".

[0091] Based on this scenario tree structure, the embodiment of the present application locates each AI training model and clarifies its applicable specific scenario path, such as "SMS data processing > number management > number validity detection". This standardized scenario description method solves the problem of inconsistent descriptions between different models, allowing the scope of application of the model to be accurately located and compared.

[0092] In addition, the embodiment of the present application will also identify the processing objectives of the model and clarify the specific problems it aims to solve, such as "detecting invalid numbers" and "identifying spam text messages". Through this description of the scene tree structure, the embodiment of the present application can clearly define the application field and processing objectives of each model, laying the foundation for subsequent accurate matching. This method is particularly suitable for processing field-specific AI models, and can accurately capture their professional characteristics and applicable boundaries.

[0093] C2: Based on the technical specifications of the initial candidate AI training model, extract technical parameters including required fields, data format and data volume restrictions to form input requirement features.

[0094] Each AI training model has specific input requirements, and only data that meets these requirements can be effectively processed. The embodiment of the present application first analyzes the technical documentation of the model, extracts information about the input data requirements, and then standardizes this information into structured input requirement features.

[0095] The embodiments of the present application identify the fields necessary for model processing, such as the number verification model may require a "mobile phone number" field, and the SMS classification model may require a "SMS content" field. The embodiments of the present application not only record the field names, but also include the semantic meanings of the fields to cope with the differences in field naming in different data sets. The embodiments of the present application clearly define the data format requirements for each required field, such as the "mobile phone number" field must be 11 digits, and the "SMS content" field should be text type and not empty. These format requirements are an important basis for evaluating the compatibility of data and models.

[0096] In addition, the embodiment of the present application will also extract the model's limit parameters on the amount of data, such as the minimum processing unit, the maximum processing capacity, the batch processing capacity, etc. For example, some models may require at least 1,000 records to be effectively trained, while other models may have significantly reduced performance when processing more than 100,000 records. At the same time, the embodiment of the present application will also identify the model's requirements for data quality, such as missing value ratio limits, outlier processing capabilities, etc. These parameters together constitute the input requirement characteristics of the model, and describe in detail the data types and conditions that the model can handle. Through this standardized input requirement description, the embodiment of the present application can accurately evaluate the compatibility of the data to be processed with each model, avoid selecting a model with mismatched input requirements, thereby improving the accuracy and effectiveness of matching.

[0097] C3: Use a standardized capability description language to process the functional characteristics of the initial candidate AI training model, break down specific functions into measurable processing units, and generate processing capability characteristics.

[0098] Different from the traditional fuzzy functional description, the embodiment of the present application adopts a standardized capability description language to convert the functional characteristics of the model into specific and measurable processing capability indicators.

[0099] The embodiments of the present application define a set of unified capability measurement dimensions, including functional dimension (specific operation types that the model can perform), performance dimension (execution efficiency and resource consumption), accuracy dimension (reliability of the results) and adaptability dimension (ability to handle edge cases), etc.

[0100] In terms of the functional dimension, the embodiment of the present application decomposes the complex functions of the model into basic processing units, such as "number format verification", "operator identification", "active status detection", etc., and records the support level of each unit (full support, partial support, or no support). In terms of the performance dimension, the embodiment of the present application quantifies indicators such as processing speed (such as the number of records processed per second), memory consumption, and parallel processing capabilities. In terms of accuracy, the embodiment of the present application records evaluation indicators such as precision, recall rate, and F1 score of the model on various tasks, as well as the confidence intervals of these indicators.

[0101] In the adaptability dimension, the embodiment of the present application evaluates the model's ability to handle non-standard situations, such as tolerance to abnormal data, missing values, and noise. These metrics together constitute the processing capability characteristics of the model, which comprehensively and specifically describes the performance of the model in various aspects.

[0102] Through this standardized capability description, the embodiments of the present application can quantitatively compare the capability differences between different models and provide an objective basis for model selection.

[0103] For example, when faced with a data set containing a large number of non-standard format numbers, the embodiment of the present application can select a model with stronger exception handling capabilities based on the score of the adaptability dimension, rather than just based on general performance indicators. This fine-grained capability description is a key link in achieving accurate model matching, which greatly improves the decision-making quality of the embodiment of the present application.

[0104] C4: Integrate the specific application field and the processing target, the input requirement characteristics and the processing capability characteristics to generate a three-dimensional feature file of the candidate AI training model as a preset model feature portrait library.

[0105] In the previous steps, the embodiments of the present application described the model features from three dimensions: applicable scenarios, input requirements, and processing capabilities. This step integrates the features of these three dimensions to form a complete model feature file.

[0106] The embodiment of the present application constructs a unified feature representation structure to ensure that the features of the three dimensions can be stored and accessed in a consistent format. This unified structure adopts a multi-level nested structure, which not only retains the independence of the features of each dimension, but also establishes the association relationship between them. The embodiment of the present application realizes the association mapping between features and clarifies the logical relationship between the features of the three dimensions.

[0107] For example, the "number validity detection" in the applicable scenario is closely related to the "number format verification" and "operator identification" in the processing capabilities. The embodiments of the present application will establish these associations for overall consideration during the matching process.

[0108] In addition, the embodiment of the present application will also calculate the integrity and reliability scores of each dimensional feature as a measure of feature quality. For features with incomplete or uncertain information, the embodiment of the present application will mark their confidence levels and adjust their weights accordingly in subsequent matching. Through this integrated processing, the embodiment of the present application ultimately generates a three-dimensional feature profile for each AI training model, which comprehensively and accurately describes all aspects of the model. These feature profiles together constitute the model feature portrait library, which provides a standardized reference basis for subsequent model matching.

[0109] Compared with traditional methods, this three-dimensional feature archive not only provides a more comprehensive model description, but also realizes the correlation representation between features, greatly improving the expressiveness of model representation and the accuracy of matching. Through this standardized model capability description system, the embodiment of the present application solves the problem of inconsistent and incomplete model description in the prior art, and provides a solid foundation for achieving accurate model matching.

[0110] It should be noted that the execution process of the feature quadruple dynamic weight matching algorithm includes: D1: Utilize the first deep learning model to calculate the semantic similarity between the structured instruction and the functional description of the candidate AI training model, and generate a natural language weight score.

[0111] The embodiment of this application uses a specially trained deep learning model as the core engine for semantic understanding. The model is based on advanced language understanding architectures, such as BERT, RoBERTa, or their domain-adapted versions, and has the ability to understand professional language in the field of SMS processing through large-scale pre-training and specific task fine-tuning.

[0112] In the matching process, the embodiment of the present application first obtains the structured instruction generated in step A, such as "Action: ValidateObject: TelephoneNumberProperty: Validity", and converts it into a natural language expression, such as "verify the validity of the phone number". At the same time, the embodiment of the present application extracts the functional description of the candidate model from the model feature portrait library, and these descriptions have been standardized, such as "detect the validity of the mobile phone number and identify invalid numbers".

[0113] Then, the embodiment of the present application uses a deep learning model to calculate the semantic similarity between the two expressions. Unlike simple vocabulary matching or vector space models, this deep learning method can understand semantic equivalence and identify semantically identical descriptions even if they are expressed in different ways.

[0114] For example, although "verify number" and "check number validity" use different words, their semantics are very similar. The embodiment of the present application will calculate this kind of semantic similarity for each candidate model and generate a natural language weight score between 0 and 1. The higher the score, the higher the semantic match between the model's functional description and user needs. This matching method based on deep semantic understanding overcomes the limitations of traditional keyword matching, can understand the similarity and equivalence at the semantic level, and significantly improves the accuracy of semantic matching. Especially for complex, multi-level task descriptions, this semantic matching method can capture deep semantics. Figure 1 Find the model that truly meets user needs.

[0115] In the embodiment of the present application, the first deep learning model adopts a hierarchical hybrid architecture design, which is specifically used to calculate the semantic similarity between structured instructions and model function descriptions. The base layer of the model adopts a pre-trained BERT-base model (12-layer Transformer encoder, hidden layer dimension 768, 12 attention heads), and is further pre-trained on a corpus in the field of SMS processing.

[0116] The domain adaptation pre-training used about 2 million text data related to SMS processing, including task descriptions, processing instructions, and result reports. The pre-training adopted a dual-task method of masked language model and next sentence prediction, with a learning rate of 2e-5, a batch size of 32, and 8 epochs. On this basis, the model added an attention enhancement layer specifically for structured instruction processing. This layer uses a multi-head cross-attention mechanism (8 attention heads) to enable the model to better capture the correspondence between structured instructions and natural language descriptions. The top layer of the model is a specific task fine-tuning layer, which contains two fully connected layers (with dimensions of 768 and 256, respectively), uses the GELU activation function, and finally calculates the output semantic matching score through cosine similarity.

[0117] The model training adopts a two-stage strategy: first, supervised learning is performed on a labeled semantic similarity dataset (about 100,000 pairs of labeled samples), using the AdamW optimizer, a learning rate of 3e-5, and a weight decay of 0.01; then online fine-tuning is performed through a feedback learning mechanism, using user feedback as a signal and a contrastive learning method to further optimize the model performance. Evaluation indicators include semantic matching accuracy, ranking quality (NDCG@5), and consistency score. On the internal test set, the model achieved a matching accuracy of 89.7%, an increase of 15.3 percentage points over the baseline model.

[0118] D2: Process the key information of the file name feature through regular expressions and rule engines, match it with the target data type of the candidate AI training model, and form a file name feature weight score.

[0119] The embodiment of the present application implements an intelligent matching mechanism based on file names. File names usually contain rich business information and can reflect the type, purpose and processing requirements of data.

[0120] The present embodiment of the application first starts with the file name features extracted in step B1, which already contain information in multiple dimensions such as time, content, status and format. Subsequently, the present embodiment of the application processes these features through regular expressions and rule engines, extracts key information from the file name and standardizes it.

[0121] For example, for a file name such as "2024Q3_SMS_NumList_Corp_FINAL.xlsx", the embodiment of the present application will recognize that it is related to key concepts such as "SMS number" and "corporate user".

[0122] Next, the embodiment of the present application matches these key information with the target data type of each candidate model in the model feature portrait library. The target data type is an important part of the model feature portrait, which describes the data category that the model is designed to process. The embodiment of the present application evaluates the degree of match between the file name features and the model target data type and calculates the matching score. This process considers not only direct matches (such as the file name explicitly contains the same keywords as the model target), but also indirect matches (such as the data type implied by the file name is compatible with the model target).

[0123] For example, the file named "Number List" has a high degree of match with the "Number Validity Detection Model", but a low degree of match with the "SMS Content Analysis Model". Finally, the embodiment of the present application generates a file name feature weight score between 0 and 1 based on the matching evaluation, reflecting the degree of fit between the processing requirements implied by the file name and the model function. This file name-based matching mechanism makes full use of the business semantics implied in file naming, providing additional decision-making basis for model selection, especially when the user's natural language description is not detailed enough, the file name feature can provide important supplementary information.

[0124] D3: Calculate the matching degree between the file content features and the candidate AI training model input requirements, and output the content feature weight score.

[0125] The embodiment of the present application obtains the file content features generated in step B2, which describe multiple aspects such as the structural organization, field characteristics, and content distribution of the data. Then, the embodiment of the present application extracts the input requirements of each candidate model from the model feature portrait library. These requirements clearly define the data conditions that the model can effectively process, including required fields, data formats, and quality requirements. Subsequently, the embodiment of the present application performs multi-level matching calculations to comprehensively evaluate the compatibility of the data content with the model input requirements.

[0126] At the field matching level, the embodiment of the present application checks whether the data contains all the necessary fields required by the model, such as the number verification model requires the "mobile phone number" field, and the SMS classification model requires the "SMS content" field.

[0127] The embodiments of the present application can handle variants and synonymous expressions of field names, such as "mobile phone number", "mobile phone number", "telephone", etc., which may point to the same semantic field.

[0128] At the format matching level, the embodiment of the present application evaluates whether the format of the data field meets the processing requirements of the model, such as whether the number is in the standard 11-digit format, whether the text is in the expected language type, etc. At the quality matching level, the embodiment of the present application analyzes the integrity, consistency and noise level of the data to evaluate whether it meets the quality requirements of the model, such as whether the proportion of missing values ​​is within an acceptable range, whether there are too many outliers, etc. At the scale matching level, the embodiment of the present application checks whether the data volume meets the processing capacity of the model, which should neither be less than the minimum processing unit of the model nor exceed the maximum processing capacity of the model.

[0129] Through this comprehensive matching analysis, the embodiment of the present application finally calculates a content feature weight score between 0 and 1, reflecting the overall compatibility between the file content and the model input requirements. This matching mechanism based on content features ensures that the selected model can effectively process the actual data, avoiding processing failures caused by data incompatibility, and is an important link in achieving accurate model matching.

[0130] D4: Analyze the data format conversion cost corresponding to the file format feature and generate a format adaptation weight score.

[0131] The embodiment of the present application obtains the file format features identified in step B3 to clarify the storage format and structural characteristics of the data. At the same time, a list of input formats supported by each candidate model is extracted from the model feature portrait library. The embodiment of the present application then performs a format compatibility analysis to evaluate the matching of the data format with the model input format.

[0132] When the data format is directly supported by the model, such as the data is in CSV format and the model accepts CSV input, the embodiment of the present application will give the highest format adaptation score because no format conversion is required.

[0133] However, in actual applications, the data format is often inconsistent with the model's expected format. In this case, the embodiment of the present application will further analyze the complexity and cost of the format conversion.

[0134] The embodiment of the present application constructs a detailed format conversion mapping table, which records the feasibility, complexity and potential risks of conversion between various formats. For example, the conversion from Excel to CSV is relatively simple and almost lossless, while the conversion from PDF table to structured data is much more complicated and may introduce errors.

[0135] Based on this conversion analysis, the embodiment of the present application calculates the format conversion cost, taking into account factors including: the technical complexity of the conversion (whether special tools or processing steps are required), the risk of data loss (the amount of information that may be lost during the conversion process), the processing time cost (the computing resources and time required for the conversion), and the reliability of the conversion results (the expected quality of the converted data).

[0136] Finally, the embodiment of the present application generates a format adaptation weight score between 0 and 1, where a higher score indicates easier format adaptation and lower cost. This format adaptability analysis ensures that the selected model can efficiently process data in a given format, avoids complex or unreliable format conversion processes, and optimizes the efficiency and reliability of the overall processing flow. In practical applications, this consideration is particularly important because inappropriate format conversion not only increases processing time, but may also introduce data errors, affecting the accuracy of the final result.

[0137] D5: Using preset weight coefficients, perform weighted calculation on the natural language weight score, file name feature weight score, content feature weight score and format adaptation weight score to obtain a comprehensive matching degree, and use the comprehensive matching degree as the final scoring result of the candidate AI training model.

[0138] In the previous steps, the embodiment of the present application evaluated the matching degree of the candidate model from four dimensions: natural language description, file name, content characteristics, and format adaptation. This step integrates the scores of these four dimensions to generate the final comprehensive matching degree.

[0139] The embodiment of the present application uses preset weight coefficients to perform weighted calculations on the scores of the four dimensions. These weight coefficients reflect the relative importance of each dimension in model matching, and the initial setting may be 40% for natural language description, 30% for file content features, 20% for file name features, and 10% for format adaptation. This weight distribution takes into account the value and reliability of the information provided by each dimension, focusing on the needs clearly expressed by the user (natural language description) and the actual characteristics of the data (content features).

[0140] Then, the embodiment of the present application performs a weighted summation calculation, adds the weighted scores of the four dimensions, and obtains a comprehensive matching degree between 0 and 1. For example, if the scores of a model in the four dimensions are 0.9 (natural language), 0.7 (file name), 0.8 (content), and 0.6 (format), its comprehensive matching degree is 0.9×40%+0.7×20%+0.8×30%+0.6×10%=0.8. This comprehensive matching degree is the final scoring result of the candidate model, reflecting the overall fit of the model with the current task requirements.

[0141] It is worth noting that although the embodiment of the present application uses preset weights for initial calculations, these weights are not fixed, but are dynamically adjusted according to user feedback in step S4. For example, if it is found that user decisions are more inclined to consider content features, the embodiment of the present application will gradually increase the weight of content features. This dynamic weight adjustment mechanism enables the embodiment of the present application to adapt to the preferences and decision-making patterns of different users and continuously optimize the accuracy of the matching algorithm. Through this multi-dimensional comprehensive scoring mechanism, the embodiment of the present application can comprehensively consider various factors, find the model that is truly most suitable for the current task, and achieve accurate intelligent matching.

[0142] Optionally, the core calculation formula of the feature quadruple dynamic weight matching algorithm in the embodiment of the present application is as follows: MatchScore = α·NL + β·FN + γ·FC + δ·FF, in: NL represents the natural language weight score, which is obtained by calculating the semantic similarity through the deep learning model: NL = cosine_similarity(E(Instruction), E(ModelDesc)), Where E() represents the semantic encoding function, which is implemented by the first deep learning model.

[0143] FN represents the file name feature weight score, which is calculated by the rule engine: FN = Σ(wi·match(ri, filename)) / Σwi, Where ri represents the rule pattern associated with the model, and wi represents the corresponding weight.

[0144] FC represents the content feature weight score, which is calculated by field matching and data compatibility: FC = λ1·FieldMatch + λ2·DataCompat, FieldMatch calculates the coverage of necessary fields, and DataCompat evaluates the compatibility of data formats.

[0145] FF represents the format adaptation weight score, which is calculated based on the format conversion cost: FF = 1 - normalize(ConversionCost), Where ConversionCost represents the computational cost of converting from the current format to the format required by the model.

[0146] The weight coefficients α, β, γ, and δ (corresponding to initial values ​​of 40%, 20%, 30%, and 10%, respectively) are dynamically adjusted based on user feedback: α' = α + η·Δα, β' = β + η·Δβ, γ' = γ + η·Δγ, δ' = δ + η·Δδ, Where η is the learning rate (usually set to 0.05), and the Δ value is calculated based on user feedback: when the user accepts the recommended model, the current weight configuration is enhanced; when the user chooses a different model, the weight is adjusted according to the difference. To ensure that the sum of the weights is 1, normalization is performed after each update: [α',β',γ',δ'] = [α',β',γ',δ'] / Σ[α',β',γ',δ'], In addition, before executing the feature quadruple dynamic weight matching algorithm, the method further includes: E1: Collect historical matching records, identify high-frequency typical subtask combinations through data mining, build a high-frequency subtask optimization pool, and output task combination patterns.

[0147] Before executing the matching algorithm, the embodiment of the present application also implements a high-frequency subtask optimization mechanism. By mining the data of historical matching records, the embodiment of the present application identifies frequently occurring typical subtask combinations, such as a task chain of "mobile phone number verification → address normalization → customer classification".

[0148] The embodiment of the present application implements an intelligent optimization mechanism based on historical experience. The embodiment of the present application continuously records the user's model matching and usage, including a complete information chain of input data features, selected models, processing results, and user feedback. These historical records constitute a valuable experience database that reflects the model application mode in actual business scenarios.

[0149] Subsequently, the present application embodiment uses advanced data mining technology to analyze these historical records and identify the frequently occurring model usage patterns. The present application embodiment adopts an improved sequence pattern mining algorithm, which not only focuses on the usage frequency of a single model, but also pays more attention to the pattern of combined use of multiple models.

[0150] For example, the embodiments of the present application may find that a processing chain such as "mobile phone number verification→address normalization→customer classification" is often used together to form a typical task combination.

[0151] For the identified high-frequency patterns, the embodiments of the present application further analyze their business background, trigger conditions, and application effects, and extract the key features and constraints in the patterns. Based on these analyses, the embodiments of the present application construct a high-frequency subtask optimization pool, store common task combination patterns in a structured manner, and mark their applicable conditions and expected effects. These task combination patterns become an important reference for subsequent model matching, especially when dealing with complex tasks. The embodiments of the present application can directly recommend verified model combination solutions without having to build a processing chain from scratch.

[0152] Through this optimization mechanism based on historical experience, the embodiments of the present application can make full use of accumulated knowledge to improve matching efficiency and accuracy. In particular, for common business scenarios, the embodiments of the present application can quickly identify the best processing solution and significantly shorten the decision-making time. This method reflects the learning ability of the embodiments of the present application. By continuously accumulating experience, the matching performance of the embodiments of the present application will continue to improve as the usage time increases.

[0153] E2: Analyze the execution process of each high-frequency subtask using the task combination pattern, extract data dependencies, execution order requirements and resource consumption constraints, and generate a set of constraint equations, including: For these high-frequency subtasks, the embodiments of the present application further analyze their execution process, extract key constraints, and formalize these constraints into a set of constraint equations. This process includes using a deep learning model to analyze historical success cases, extract data flow rules and resource constraints, and convert constraints into a computational model.

[0154] E2.1: Call the first deep learning model to process the model combination pattern in the historical successful cases, and output the model link prediction strategy through pattern recognition and feature learning.

[0155] The embodiments of the present application implement an advanced pattern learning and prediction mechanism.

[0156] The embodiment of the present application selects successful cases with good processing effects from historical records, which contain complete data features, model selection and processing result information. Then, the embodiment of the present application calls a specially designed deep learning model to analyze these successful cases. The model adopts a composite architecture that combines sequence modeling capabilities (such as LSTM or Transformer) and graph structure learning capabilities (such as neural networks in the figure), which can simultaneously capture the temporal patterns and interdependencies used by the model.

[0157] In the learning process, the embodiment of the present application first represents each successful case as a structured feature sequence, including data features, task requirements, context information, etc. Then, the deep learning model learns the mapping relationship between input features and the optimal model combination from these cases through self-supervised learning.

[0158] Through training on a large number of cases, the model gradually mastered the implicit rules of model selection in different scenarios and was able to predict the most likely successful model combination based on new input features.

[0159] After learning is completed, the embodiment of the present application outputs a model link prediction strategy, which is a decision-making embodiment of the present application that can quickly generate model combination suggestions based on input features. This strategy not only considers the applicability of a single model, but also considers the coordination relationship and sequence dependency between models, and can recommend an overall optimal processing flow.

[0160] For example, when faced with SMS number data containing multiple abnormal situations, the strategy may recommend using the "data cleaning model" to process the abnormal format first, then using the "number validity verification model" to detect the validity, and finally using the "number classification model" for classification statistics. This prediction strategy based on deep learning greatly improves the ability of the embodiment of the present application to handle complex tasks, and can automatically build the optimal model processing chain based on historical experience, reducing the complexity of manually designed processing procedures.

[0161] The neural routing solver in the embodiment of the present application is built based on a graph neural network architecture and includes three key components: a feature encoder, a routing predictor, and a search strategy generator.

[0162] The feature encoder uses a multi-layer perceptron structure, including 3 hidden layers (dimensions are 256, 512, and 256 respectively), and uses the ReLU activation function to map the input features into high-dimensional representations. The specific implementation is: h1= ReLU(W1x+b1), h2= ReLU(W2h1+b2), h3= ReLU(W3h2+b3), Where x is the input feature, W and b are the weight matrix and bias vector respectively.

[0163] The route predictor is based on the Transformer architecture, which includes 4 attention heads and 2 encoder layers, and can capture the complex correlation patterns between features and models. The attention mechanism calculation formula is: Attention(Q,K,V) = softmax(QK T / dk)V Where Q, K, V are query, key-value and value matrices respectively, and dk is the dimension of the key-value vector.

[0164] The search strategy generator uses a policy gradient network, and the output layer uses a Softmax function to generate the probability distribution of different search paths: π(a|s) = softmax(Wπhπ + bπ), Where s is the current state, a is a possible action, and hπ is the hidden representation of the policy network.

[0165] E2.2: Analyze the high-frequency subtasks according to the model link prediction strategy, extract data flow rules and computing resource limitation constraint elements, and form formalized constraint conditions.

[0166] The embodiment of the present application conducts an in-depth analysis of the identified high-frequency subtasks based on the model link prediction strategy outputted in the previous step. The focus of the analysis is to understand the dependencies and execution conditions between tasks and extract constraint elements that can be formally expressed.

[0167] In terms of data flow, the embodiments of the present application analyze the data transfer rules between models, including the necessary input-output matching relationship, data format conversion requirements and the dependencies of intermediate results. For example, the embodiments of the present application may identify that the "SMS content analysis model" needs to receive the valid records filtered by the "SMS number verification model" as input, which constitutes a data flow constraint.

[0168] In terms of computing resources, the embodiments of the present application evaluate the resource requirements and performance characteristics of each model, including processing time, memory consumption, parallelism, etc. These resource constraints are crucial to evaluating the feasibility of the entire processing chain, especially in large-scale data processing scenarios.

[0169] In addition, the embodiments of the present application also analyze business logic constraints, such as rules that certain models must be executed under specific conditions or must be executed in a specific order.

[0170] For example, "Customer Information Matching" must be performed after "Number Validity Verification" to avoid unnecessary matching operations on invalid numbers.

[0171] Through this comprehensive constraint analysis, the embodiments of the present application convert various implicit dependencies and constraints into clear formal constraint expressions. These constraint expressions use standardized formats, such as "ModelAoutput->ModelBinput" for data flow constraints, "ModelC.time<5min" for time constraints, etc. These formal constraints provide clear evaluation criteria for subsequent model search and optimization, ensuring that the model combination recommended by the embodiments of the present application is not only functionally matched, but also feasible under actual operating conditions.

[0172] E2.3: Convert the constraints into a computational model and establish a quantifiable set of constraint equations to guide subsequent model search and optimization processes.

[0173] The embodiment of the present application realizes an advanced constraint modeling and quantitative calculation mechanism.

[0174] The embodiment of the present application converts the formalized constraint conditions extracted in the previous step into a strict mathematical representation to construct a computable constraint equation set. This conversion process uses a hybrid constraint solving technique to uniformly represent different types of constraints as a computational model.

[0175] For data flow constraints, the embodiment of the present application establishes a model dependency representation based on a directed graph, where each model is regarded as a node in the graph, the data flow relationship is regarded as a directed edge, and a compatibility weight is assigned to each edge to indicate the degree of matching of data transmission.

[0176] For resource constraints, the embodiment of the present application establishes a representation based on linear inequalities, represents resource requirements such as processing time and memory consumption as a linear combination of model selection variables, and sets the corresponding upper threshold. For business logic constraints, the embodiment of the present application uses predicate logic expressions for formalization, converting relationships such as conditional execution and sequential dependencies into logical constraints. After normalization, these constraint equations constitute a unified set of constraint equations that describe all the conditions that must be met for an effective model combination.

[0177] The embodiment of the present application also designs a solution priority for the constraint equation group, distinguishing necessary constraints (such as data compatibility) from secondary constraints (such as performance optimization) to ensure that reasonable trade-offs can be made when constraints conflict. The final constructed constraint equation group not only describes the boundaries of feasible solutions, but also contains optimization objective functions, such as minimizing total processing time or maximizing result accuracy. This quantifiable constraint model provides clear evaluation criteria for subsequent model searches, can quickly judge the feasibility and pros and cons of candidate solutions, and greatly improves search efficiency. By converting business requirements and technical limitations into precise mathematical models, the embodiment of the present application achieves accurate expression and efficient processing of complex constraints, laying the foundation for optimizing model combinations.

[0178] E3: Process the search space of candidate AI training models according to the set of constraint equations, perform pruning operations, and form an optimized set of candidate AI models.

[0179] Based on these constraint equations, the embodiment of the present application performs pruning operations on the search space of candidate models to exclude model combinations that obviously do not meet the conditions, thereby improving search efficiency. This optimization mechanism is particularly suitable for complex task scenarios that require multiple models to be processed collaboratively.

[0180] In order to further optimize the search efficiency in the model matching process, the embodiment of the present application implements a high-frequency subtask constraint optimization system. First, by performing data mining on historical matching records, the embodiment of the present application identifies typical subtask combinations that often appear in the processing process, such as the task chain of "mobile phone number verification → address normalization → customer classification". These high-frequency subtask combinations are stored in the optimization pool, and a standardized feature description is established for each combination.

[0181] Secondly, the embodiment of the present application implements the function of automatic extraction of constraints. For each high-frequency subtask, the embodiment of the present application analyzes the key constraints in its execution process, including data dependencies, execution order requirements, resource consumption restrictions, etc. These constraints are formalized as a computable set of constraint equations. For example, for the task of "mobile phone number verification", the embodiment of the present application will extract specific constraints such as "the input must be 11 digits" and "the processing delay does not exceed 100ms". These constraint equations are used to narrow the subsequent search space.

[0182] The embodiments of the present application implement an efficient search space optimization mechanism.

[0183] In practical applications, the number of possible model combinations grows exponentially with the growth of the model library size, and it is unrealistic to conduct a detailed evaluation of all combinations. To solve this problem, the embodiment of the present application uses the previously constructed constraint equation group to perform intelligent pruning of the search space.

[0184] The embodiment of the present application performs a preliminary screening of all candidate models to exclude models that clearly do not meet the basic requirements, such as models whose input formats are completely incompatible or lack necessary functions. This initial screening process significantly narrows the basic search space. Subsequently, the embodiment of the present application applies a constraint propagation algorithm to perform deep pruning based on hard constraints in the constraint equation group (such as data flow relationships and necessary conditions).

[0185] For example, if a constraint requires that the output of model A must be able to serve as the input of model B, the embodiment of the present application will exclude all variants of model A whose output is incompatible with model B, thereby greatly reducing the number of combinations that need to be considered.

[0186] For complex model chains, the embodiment of the present application adopts a dynamic programming method to gradually build a solution space starting from simple sub-problems, avoiding repeated exploration of infeasible substructures.

[0187] On this basis, the embodiment of the present application also applies a heuristic search strategy to preferentially explore model combination patterns with high historical success rates to further improve search efficiency.

[0188] Through this series of optimization operations, the embodiment of the present application ultimately forms a significantly reduced set of candidate AI models, and these model combinations all meet the basic requirements of the constraint equation group and have high feasibility.

[0189] For example, for an embodiment of the present application that contains hundreds of basic models, there are tens of thousands of possible combinations, but after constrained pruning, only dozens of high-quality candidate combinations may need to be considered. This search space optimization not only greatly improves the speed of model matching, but also improves the matching quality by eliminating unreasonable combinations, enabling the embodiment of the present application to quickly recommend the optimal model processing solution in complex scenarios.

[0190] A feature compression mechanism is adopted in the feature quadruple dynamic weight matching algorithm, including: F1: Obtain the original feature space data, construct a feature pyramid structure through multi-level feature extraction, and output feature combinations at different granularity levels.

[0191] The feature pyramid structure in the embodiment of the present application adopts a multi-level design, usually including 3 to 5 levels. At the bottom layer (1st layer), the complete original features are retained, the dimensions are usually between 200-500, and all fine-grained information such as specific field names, data formats, distribution characteristics, etc. are included. In the middle layer (2nd-3rd layer), the dimension is reduced to 50-100 through principal component analysis (PCA) and feature aggregation, and the main feature patterns are retained. At the top layer (4th-5th layer), it is further compressed to a highly abstract feature of 10-20 dimensions, containing only core information such as data types and main tasks.

[0192] A nonlinear mapping relationship is used between the layers, and an autoencoder network is used to achieve feature compression and reconstruction. Specifically, the compression from the i-th layer to the i+1-th layer uses the encoder network Ei, and the reconstruction from the i+1-th layer to the i-th layer uses the decoder network Di. Both the encoder and decoder use a fully connected neural network structure, with a BatchNormalization layer and a LeakyReLU activation function in the middle to ensure the nonlinear expression ability of feature conversion. These networks are trained by minimizing the reconstruction error ||x - Di(Ei(x))||² to ensure that the compression process retains the most informative features.

[0193] The embodiments of the present application implement an innovative multi-scale feature representation mechanism.

[0194] In the model matching process, although the full amount of original features contains complete information, the computational complexity is high and they are easily affected by noise.

[0195] To solve this problem, the embodiment of the present application adopts a feature pyramid structure to achieve feature representation at different abstract levels.

[0196] The embodiment of the present application obtains the complete feature space data extracted in step S1, including all original information such as natural language features, file name features, content features, and format features. Then, the embodiment of the present application constructs a feature pyramid from bottom-level details to top-level summaries through a multi-level feature extraction algorithm.

[0197] At the bottom of the pyramid, the embodiment of the present application retains complete original features, such as detailed field names, specific data distribution, and complete file format information. The feature representation of this layer is the most detailed, but also the most complex.

[0198] At the middle layer, the embodiment of the present application generates a more abstract feature representation through feature aggregation and dimensionality reduction operations. For example, the detailed field information is aggregated into a summary feature such as "contains telephone number information", and the specific data distribution is simplified into an evaluation result such as "good data integrity". These middle-layer features retain the main characteristics of the original data while significantly reducing the feature dimension.

[0199] At the top level of the pyramid, the embodiment of the present application generates highly abstract feature summaries, such as core feature descriptions such as "SMS number data" and "validity verification required". Although these top-level features have a small amount of information, they capture the essential characteristics of the task and are suitable for fast matching.

[0200] Through this pyramid structure, the embodiment of the present application represents the features of the same data at different granularity levels, which is conducive to rapid screening while retaining the possibility of detailed verification. The features of each layer in the feature pyramid maintain a clear hierarchical association. The upper layer features are an effective summary of the lower layer features, and the lower layer features are a detailed expansion of the upper layer features. This multi-level feature representation greatly improves the flexibility and efficiency of feature processing and provides a solid foundation for subsequent iterative matching.

[0201] F2: Based on the feature combination, perform a top-down iterative matching operation, use the compressed feature data to screen candidate AI training models, and generate preliminary matching results, including: F2.1: Perform preliminary screening through the highest level compressed features of the feature pyramid structure, determine the category range of the candidate AI training models, and output the model category data.

[0202] Before full matching, the embodiment of the present application first uses the highly compressed features at the top level of the feature pyramid to perform preliminary screening to quickly determine potentially relevant model categories.

[0203] This process is similar to the way human experts think, first determining the general direction based on the general characteristics of the task, and then conducting detailed analysis.

[0204] Specifically, the embodiment of the present application obtains the compressed features of the top layer of the feature pyramid, which highly summarize the core properties of the data and tasks, such as basic labels such as "SMS data", "number verification", and "content classification". Then, the embodiment of the present application quickly matches these compressed features with the top-level features of each model in the model feature portrait library. Since the top-level features are low in dimension and highly abstract, this matching process has a small amount of calculation and can be completed quickly. Based on the matching results, the embodiment of the present application calculates the preliminary relevance of each model category to the current task.

[0205] The "model category" here refers to a collection of models with similar functions, such as "number verification model", "SMS classification model", etc.

[0206] The embodiment of the present application selects several model categories with higher scores as candidate sets based on the relevance score. For example, for the task of "checking the validity of SMS numbers", the embodiment of the present application may give priority to "number verification" and "data verification" models, while excluding "content analysis" and "customer grouping" models. This preliminary screening greatly narrows the scope of models that need to be examined in detail and improves the efficiency of subsequent matching.

[0207] At the same time, the embodiment of the present application records the preliminary matching evidence and credibility score of each candidate category as a reference for subsequent fine matching. Through this top-down screening strategy, the embodiment of the present application can quickly focus on the most likely relevant model subset, avoiding invalid calculations of obviously irrelevant models, and significantly improving matching efficiency, especially when the model library is large.

[0208] F2.2: Using the model category data, perform feature matching refinement at each level of the feature pyramid in turn, determine the candidate AI training model through layer-by-layer screening, and generate the screening results.

[0209] After determining the preliminary model category range, the embodiment of the present application begins to perform a gradually refined matching analysis at each level of the feature pyramid. This hierarchical matching strategy is similar to the "progressive approximation" method in scientific research, which continuously increases the accuracy and depth of the analysis and ultimately determines the best solution.

[0210] The embodiment of the present application performs a second round of matching in the middle layer of the feature pyramid. The middle layer features are more specific than the top layer, but still maintain a moderate level of abstraction, such as "a data set containing standard format phone numbers", "double verification of validity and location is required", and other more detailed task descriptions.

[0211] The embodiment of the present application matches the middle-layer features with each specific model in the candidate model category screened out in the previous step to further narrow the candidate range. For example, in the "number verification type" model, some models may focus on format verification, while others may include the function of identifying the place of origin. The embodiment of the present application will distinguish them based on the matching results of the middle-layer features. Subsequently, the embodiment of the present application enters the bottom layer of the feature pyramid and uses the most detailed original features for the final match. The bottom-level features contain all the detailed information, such as specific field names, data format requirements, exception handling capabilities, etc.

[0212] The embodiment of the present application performs a comprehensive and detailed evaluation of the remaining high-potential candidate models, taking into account all relevant factors, and generates a final matching score. Through this top-down, layer-by-layer refinement matching process, the embodiment of the present application effectively combines the advantages of fast screening and accurate matching. In each layer of matching, the embodiment of the present application will exclude models that obviously do not meet the requirements, so that the number of models that need to be evaluated in detail is reduced layer by layer, greatly improving the matching efficiency.

[0213] Finally, the embodiment of the present application generates a screening result containing a small number of high-quality candidate models, which have passed multi-level screening and have a high matching credibility. Compared with the traditional one-time full-volume matching, this hierarchical matching strategy not only improves efficiency but also ensures matching quality, and is particularly suitable for processing complex tasks and large-scale model libraries.

[0214] F2.3: Evaluate and verify the screening results, store the feature weight combinations and matching strategies that meet the preset thresholds into the algorithm library to form a reusable feature matching solution.

[0215] The embodiment of the present application realizes an advanced algorithm learning and knowledge accumulation mechanism. After completing the multi-level screening, the embodiment of the present application not only focuses on the final matching results, but also pays more attention to the effectiveness evaluation and experience extraction of the entire screening process. The embodiment of the present application conducts a comprehensive quality assessment of the screening results, using a multi-dimensional evaluation index system. These indicators include accuracy indicators (the overlap between the screening results and the optimal selection), efficiency indicators (the computational cost and time consumption of the screening process), stability indicators (the sensitivity of the screening results to input changes) and interpretability indicators (whether the basis of the screening decision is clear), etc. The embodiment of the present application will combine historical verification data and expert knowledge base to set reasonable evaluation criteria and weights for each evaluation indicator.

[0216] During the evaluation process, the embodiment of the present application will focus on analyzing the contribution of feature weight combinations. Feature weight combinations refer to the weight configuration of different feature dimensions in the matching algorithm, such as giving a higher weight to the field format feature and a lower weight to the file name feature.

[0217] The embodiment of the present application analyzes the impact of different weight combinations on the matching results through the control variable method, and identifies the weight configuration mode that is particularly effective for the current task type. At the same time, the embodiment of the present application also evaluates the effectiveness of the entire matching strategy, including the construction method of the feature pyramid, the setting of decision points for hierarchical screening, and the pruning strategy.

[0218] For those feature weight combinations and matching strategies whose evaluation results are significantly better than the benchmark level and reach the preset quality threshold, the embodiment of the present application will standardize them and store them in the algorithm library. This storage process includes not only the record of weight parameters and strategy rules, but also the clear definition of usage conditions and applicable scenarios.

[0219] For example, the embodiment of the present application may identify that a specific weight combination is particularly effective in processing the task of "short message data cleaning containing a large number of invalid numbers", and this combination will be saved together with its applicable conditions. These conditions are usually defined based on data feature patterns (such as data types, structural characteristics, abnormal patterns, etc.) and task semantic features (such as operation types, processing targets, etc.).

[0220] Optionally, in the process of performing top-down iterative matching, the embodiment of the present application adopts an advanced neural routing solver to optimize the model search path. The neural routing solver is built on an improved graph neural network architecture and includes three key components: a feature encoder, a routing predictor, and a search strategy generator.

[0221] The feature encoder adopts a multi-layer perceptron structure, which contains 3 hidden layers (with dimensions of 256, 512, and 256 respectively). It uses the ReLU activation function to map the input features into high-dimensional representations.

[0222] The route predictor is based on the Transformer architecture, which consists of 4 attention heads and 2 encoder layers, and can capture complex correlation patterns between features and models.

[0223] The search strategy generator uses a policy gradient network, and the output layer uses a Softmax function to generate the probability distribution of different search paths. The training of the neural routing solver adopts a combination of supervised learning and reinforcement learning, using historical successful matching cases as supervision signals, and optimizing the search efficiency through a reward function, which is designed as a weighted combination of matching accuracy and search steps. In this way, the neural routing solver can intelligently guide the search process, significantly reduce the search space, and improve matching efficiency.

[0224] In addition, the embodiment of the present application will also establish evaluation records for the stored matching solutions, including historical usage, success rate statistics, and applicable case analysis. These records are continuously updated as the solutions are continuously used, forming a dynamic performance evaluation mechanism. As time goes by and data accumulates, the embodiment of the present application can more and more accurately evaluate the actual effect of each matching solution and adjust its priority in future tasks accordingly.

[0225] By establishing such an algorithm library that is constantly learning and optimizing, the embodiments of the present application realize the accumulation and reuse of matching experience. When faced with a new matching task, the embodiments of the present application first check whether there are applicable stored solutions. If a solution with a high degree of matching is found, it can be directly applied, greatly improving decision-making efficiency. Even if a completely matching solution cannot be found, the embodiments of the present application can also refer to the successful experience of similar scenarios and make reasonable strategy adjustments. This learning mechanism based on experience accumulation enables the embodiments of the present application to continuously improve their own performance, especially for common task types, the matching accuracy and efficiency will be significantly improved with the increase in the number of uses.

[0226] Ultimately, this step establishes a self-evolving feature matching knowledge base, making the embodiment of the present application not just a static matching tool, but an intelligent embodiment of the present application that can continuously learn and optimize. This mechanism is particularly suitable for enterprise-level application scenarios, and can adapt to the business model and data characteristics of a specific organization, providing increasingly accurate model matching services. By continuously accumulating and optimizing matching strategies, the embodiment of the present application can continuously improve processing efficiency while maintaining a high matching accuracy rate, providing users with an increasingly smooth intelligent matching experience.

[0227] F3: Based on the preliminary matching results, a feature weight adjustment algorithm is constructed to dynamically optimize the feature matching process and output an optimized weight matching algorithm.

[0228] After completing the preliminary multi-level matching, the embodiment of the present application analyzes the key decision points and influencing factors in the matching process and identifies the links that may have room for optimization. The embodiment of the present application evaluates the actual contribution of different features in the current task and identifies which features play a decisive role in the correct matching and which features are relatively minor.

[0229] For example, the embodiment of the present application may find that in the number verification task, the data field format feature is more critical than the file name feature. Based on these analyses, the embodiment of the present application dynamically constructs a feature weight adjustment algorithm to optimize the feature weight allocation for the current task type. This weight adjustment is context-aware and will change dynamically according to the task characteristics and data characteristics, rather than a simple fixed rule.

[0230] In addition, the embodiment of the present application also analyzes the key decision paths in the matching process to identify which matching steps have the greatest impact on the final result.

[0231] For example, for some tasks, the initial screening at the top of the feature pyramid may be highly accurate, while for other tasks, it may be necessary to rely on detailed features at the bottom to make a correct judgment. Based on these patterns, the embodiments of the present application optimize the selection strategy of the feature hierarchy to customize the most efficient matching path for different types of tasks.

[0232] At the same time, the embodiment of the present application also learns effective feature combination patterns from the preliminary matching results to identify which feature combinations are particularly effective for specific types of tasks. For example, for SMS number verification tasks, a feature combination such as "field format + data distribution + model accuracy" may be particularly discriminative.

[0233] The embodiments of the present application encode these effective combinations into optimized matching rules to further improve matching efficiency. Through this series of adaptive optimizations, the embodiments of the present application output highly customized weight matching algorithms for the current task, which can more accurately and efficiently identify the most suitable model. This dynamic optimization mechanism enables the embodiments of the present application to continuously learn and improve, and always maintain the best matching performance for different types of tasks.

[0234] In one embodiment, an adaptive heuristic algorithm generator is designed to dynamically generate optimization strategies for new matching problems. The generator consists of three parts: a pattern recognition engine, a strategy synthesizer, and a performance evaluator.

[0235] The pattern recognition engine adopts a densely connected convolutional neural network architecture, which consists of 5 convolution layers (the convolution kernel sizes are 3×3, 5×5, 3×3, 3×3, and 1×1 respectively) and 3 pooling layers, which can identify typical problem patterns and challenge points from the matching history.

[0236] The strategy synthesizer is implemented based on program synthesis technology, using a predefined basic algorithm component library (including 15 basic sorting algorithms, 8 filtering methods and 12 feature transformation functions) as building blocks, and automatically combines them to generate new heuristic algorithms through genetic programming methods.

[0237] The performance evaluator uses an A / B testing framework to compare and evaluate the newly generated algorithm with the benchmark algorithm on historical data, and calculate accuracy improvement, efficiency improvement and stability indicators.

[0238] When the performance of the newly generated heuristic algorithm exceeds the preset threshold (usually requiring at least a 5% increase in accuracy or a 30% increase in efficiency), the algorithm will be added to the algorithm library and assigned an applicable scenario label. These dynamically generated heuristic algorithms can provide optimized solutions for specific types of matching problems, such as processing high-noise SMS data, parsing complex nested file structures, or processing multi-language mixed content, greatly enhancing the adaptability and performance of the embodiments of the present application.

[0239] S3: Based on the candidate AI training model, an intelligent decision-making assistance interface is constructed to analyze and display the processing logic flow chart, feature analysis report and historical usage records of the candidate AI training model to generate user feedback results.

[0240] This step implements a human-machine collaborative verification mechanism, transforming the matching engine's technical capabilities into a usable interactive experience.

[0241] The embodiment of this application designs a clear and intuitive three-pane comparison view interface: the left pane displays the processing logic flow chart of the recommended model, using a visual way to present the working principle and processing steps of the model; the middle pane displays the feature analysis report of the data to be processed, including key information such as data structure, distribution characteristics, and abnormal points; the right pane provides the historical use record of the model, such as usage frequency, success rate, typical application scenarios, etc. This multi-dimensional information display method takes into account the cognitive needs of user decision-making, so that users with non-technical backgrounds can also understand the reasons for model selection and expected effects.

[0242] For example, when the embodiment of the present application recommends a "SMS number validity verification model", the user can understand through the flowchart the processing flow of the model that first checks the number format and then verifies the operator's validity; view the format distribution and possible anomalies of the number in the current data through the feature analysis report; and understand the accuracy and processing speed of the model in similar tasks through historical records.

[0243] Based on this information, the user can make a decision to confirm or adjust the recommendation of the embodiment of the present application and provide corresponding feedback. This auxiliary decision interface greatly reduces the professional threshold for model selection, allowing business personnel to complete model selection independently without relying on technical experts. At the same time, the embodiment of the present application records the user's decision-making process and feedback, providing a data basis for subsequent optimization.

[0244] In this step, the embodiment of the present application designs an intelligent decision-making support interface, which adopts a three-pane comparison view layout. Figure 4As shown in the figure, the left pane displays the processing logic flow chart of the recommendation model, which intuitively shows the working principle of the model; the middle pane displays the feature analysis report of the data to be processed to help users understand the characteristics of the data; the right pane provides the historical usage record of the model, including usage frequency, success rate and other information.

[0245] This visual display design fully considers the cognitive needs of users' decision-making, and reduces the cognitive burden of users and improves decision-making efficiency by providing multi-dimensional decision support information. Users can confirm or adjust the model recommended by the embodiment of this application based on this information and provide feedback.

[0246] For example, when the embodiment of the present application recommends the "SMS number validity verification model", users can understand the processing flow of the model through the interface (such as checking the number format first, then verifying the operator validity), view data feature analysis (such as number format distribution, possible anomalies), and historical application (such as accuracy in similar tasks). This information helps users make more informed decisions.

[0247] When implementing the intelligent decision-making support interface, the embodiment of the present application adopts a visualization architecture based on Web components. The front-end visualization engine is built based on D3.js and ECharts libraries, and adopts a componentized design mode to decompose the interface into reusable view components.

[0248] The core view components include the model flowchart renderer, data feature visualizer, and history analyzer. The model flowchart renderer uses a directed acyclic graph layout algorithm (specifically adopting a hierarchical layout strategy with automatic node spacing adjustment), and uses SVG vector graphics technology to achieve highly interactive process display, supporting operations such as zooming, node expansion, and path highlighting.

[0249] The data feature visualizer integrates multiple chart types (including heat maps, radar charts, and parallel coordinates charts), and can automatically select the most suitable visualization method based on the data characteristics, such as using bar charts for categorical features and box plots for numerical distribution.

[0250] The History Analyzer uses time series visualization technology to display the changing trends of model application frequency, success rate and user satisfaction.

[0251] The interface communicates with the backend through RESTful API, uses JSON format to transmit data, and the frontend implements a data caching mechanism to improve response speed.

[0252] In terms of user experience design, the interface follows the principle of progressive information disclosure, initially displaying only key decision-making information, and users can expand detailed content as needed.

[0253] The interaction design adopts a consistent feedback mechanism. All user operations have clear visual and functional feedback, and the perceptibility of state changes is enhanced through animated transitions. In terms of accessibility design, the interface supports keyboard navigation and screen readers, and meets the WCAG2.1AA level standard. This highly interactive and information-rich decision-making interface significantly reduces the user's cognitive burden and makes complex model matching decisions intuitive and controllable.

[0254] S4: Utilize the user feedback results to update the weight parameters in the feature quadruple dynamic weight matching algorithm, re-evaluate and classify the AI ​​training model in the preset model feature portrait library according to the updated weight parameters, and generate an optimized model feature portrait library.

[0255] This step establishes a complete feedback learning closed-loop mechanism and realizes the self-evolution capability of the embodiment of the present application. After the user confirms or modifies the recommendation result of the embodiment of the present application, the embodiment of the present application will record the difference between the manual decision and the suggestion of the embodiment of the present application, and analyze the characteristic patterns of these differences. For example, if the user selects a model different from the one recommended by the embodiment of the present application when processing SMS number data multiple times, the embodiment of the present application will analyze the common characteristics of these situations, such as whether they are related to specific data formats or content features.

[0256] Based on these analyses, the embodiment of the present application automatically adjusts the weight parameters in the dynamic weight matching algorithm of the feature quadruple. If it is found that the user's decision is more dependent on the content features, the embodiment of the present application may increase the proportion of the content feature weight score; if it is found that the natural language description is not enough to accurately express the user's intention, the embodiment of the present application may reduce the proportion of the natural language weight score. This dynamic adjustment mechanism enables the embodiment of the present application to gradually adapt to the preferences and decision-making patterns of specific users or organizations, and improve the personalization and accuracy of matching.

[0257] At the same time, the embodiment of the present application will also update and optimize the model feature portrait library based on user feedback. If a model is frequently used for a specific type of task, the embodiment of the present application will strengthen the relevance of the model in the corresponding scenario; if the actual performance of a model deviates from its feature description, the embodiment of the present application will adjust its feature description to make it more accurate.

[0258] For example, if a model that claims to be able to process numbers in multiple formats performs poorly for certain formats in actual use, the embodiment of the present application will modify its processing capability description accordingly. This dynamic update of model features enables the feature profile library to be continuously improved and more accurately reflect the true capabilities of the model.

[0259] Through this closed-loop optimization mechanism, the embodiment of the present application achieves the self-evolution capability of "more accurate with use", overcoming the limitation of static matching strategies in the prior art that cannot adapt. This mechanism is particularly suitable for long-term enterprise-level application scenarios, and can continuously improve the performance of the embodiment of the present application as the usage increases, forming a positive feedback loop. Actual tests show that after about 100 user feedbacks, the matching accuracy of the embodiment of the present application can be improved by 15-20 percentage points, greatly reducing the need for manual intervention.

[0260] This step is one of the key innovations of the entire embodiment of the present application, and realizes the self-evolution capability of the embodiment of the present application. When the user confirms or modifies the recommendation result of the embodiment of the present application, the embodiment of the present application will record the difference between the manual decision and the suggestion of the embodiment of the present application, and adjust the weight parameters in the dynamic weight matching algorithm of the feature quadruple accordingly.

[0261] For example, if a user selects a model different from the one recommended by the embodiment of the present application multiple times in a specific scenario, the embodiment of the present application will analyze the characteristic pattern of the difference and adjust the weight coefficient accordingly. For example, the weight of the file content feature may be increased, or the weight of the file name feature may be reduced. This dynamic adjustment mechanism enables the embodiment of the present application to continuously learn the user's preferences and decision-making patterns, and gradually improve the matching accuracy.

[0262] At the same time, the embodiment of the present application will also update and optimize the model feature portrait library based on user feedback. If a model is frequently used for a specific type of task, the embodiment of the present application will strengthen the relevance of the model in the corresponding scenario; if the actual performance of a model deviates from its feature description, the embodiment of the present application will adjust its feature description to make it more accurate.

[0263] This closed-loop optimization mechanism solves the problem that static matching strategies cannot evolve on their own, allowing the embodiments of the present application to continuously improve performance as they are used, forming a virtuous circle.

[0264] For example, a telecommunications company needs to process a batch of SMS marketing data, including customer mobile phone numbers, sent content, and status records. The user uploaded an Excel file named "2023Q4_SMS Marketing Activity_Result Statistics_V3.xlsx" and provided a natural language description: "It is necessary to analyze this batch of SMS sending records, identify invalid numbers and generate detailed reports, and at the same time count the conversion rates of various types of SMS."

[0265] In the feature acquisition stage, the embodiment of the present application first analyzes the natural language description. By calling the pre-trained scenario-based intent recognition model, the description is divided into four semantic blocks: "analyze SMS sending records", "identify invalid numbers", "generate detailed reports" and "count conversion rates". From these semantic blocks, the embodiment of the present application extracts core action instructions ("analysis", "identification", "generation", "statistics"), target objects ("SMS sending records", "invalid numbers", "detailed reports", "conversion rates") and other elements to form preliminary semantic information. Then, the embodiment of the present application maps these semantic elements to the preset intent knowledge graph to form standardized structured instructions: "Action:Analyze+Identify+GenerateObject: SMS_Records+Invalid_Numbers+ReportProperty:Conversion_Rate".

[0266] At the same time, the embodiment of the present application performs multi-dimensional feature extraction on the uploaded files. The file name is analyzed by the dynamic rule engine to extract time information (fourth quarter of 2023), content information (SMS marketing campaign results statistics), status information (V3 version) and format information (Excel). The file content is detected to identify key fields such as "mobile phone number", "SMS content", "sending status", "customer feedback", etc., and the data format and distribution characteristics are determined by random sampling, such as the mobile phone number is in 11-digit format, the average length of the SMS content is 70 characters, and the sending status contains two values ​​of "success" / "failure".

[0267] In the model matching stage, the embodiment of the present application calls the feature quadruple dynamic weight matching algorithm. First, the system calculates the semantic similarity between the structured instructions and the candidate model function description, and the "SMS Analysis and Report Generation Model" obtains a semantic similarity score of 0.88. Then, the file name features are processed by the rule engine to determine that the match with the SMS marketing analysis task is 0.85. The file content features are matched with the model input requirements, and a content matching score of 0.92 is calculated. The file format features are analyzed to determine that the format is fully compatible, with a score of 1.0. Through weighted calculation (semantics 40%, file name 20%, content 30%, format 10%), the comprehensive matching degree of the model is 0.895.

[0268] In the decision interface presentation stage, the system built a three-pane comparison view interface. The left pane shows the processing flow chart of the "SMS Analysis and Report Generation Model", including four main steps: data preprocessing, number validity detection, classification statistics, and report generation. The middle pane shows the feature analysis report of the uploaded data, highlighting that about 8% of the records may contain invalid numbers and marking several major types of SMS. The right pane provides the historical use record of the model, showing an average accuracy of 93.5% on similar tasks and an average processing time of 2.5 minutes per 10,000 records.

[0269] The user confirmed the selection of the model through the interface, and fine-tuned the invalid number types that were automatically detected, adding the detection requirement for the specific category of "unregistered numbers". The system recorded this feedback and used it to update the weight matching algorithm. Specifically, the system increased the weight of SMS number field format recognition and added the "support unregistered number recognition" capability tag to the "SMS analysis and report generation model" in the model feature portrait library.

[0270] Through this complete process, the system successfully matches the user's data processing needs with the most appropriate AI model, and further optimizes the matching algorithm through user feedback, reflecting the practicality and effectiveness of the technical solution of this application.

[0271] like Figure 5 As shown, the present application also provides an SMS AI training model intelligent matching system, including a feature acquisition module, a model matching module, a decision interface module and an optimization update module, which respectively implement each step in the above method to form a complete technical implementation architecture.

[0272] To summarize, this application comprehensively solves the main problems in the prior art by constructing a multi-dimensional input parsing system, implementing a dynamic weight matching algorithm, designing a standardized model feature portrait system, and establishing a feedback learning closed-loop mechanism, achieving intelligent and precise matching of SMS AI training models, and greatly improving data processing efficiency and accuracy.

[0273] For example, when a business person with a non-technical background needs to process a batch of customer SMS number data, he only needs to upload the data file and simply describe "need to check whether this batch of numbers is valid", and the system can automatically analyze the file features, understand the user's intention, and recommend the most appropriate "number validity verification model". Through the intuitive decision-making interface, users can quickly confirm or adjust system recommendations without having to understand the complex details of the AI ​​model. This greatly reduces the threshold for use and improves work efficiency.

[0274] The present disclosure also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the SMS AI training model intelligent matching method described in the above method embodiment are executed. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0275] In addition, an embodiment of the present disclosure also provides a computer program product, which stores a computer program. When the computer program is executed by a processor, the steps of the SMS AI training model intelligent matching method provided in any of the above embodiments of the present disclosure are executed. For details, please refer to the above method embodiments, which will not be repeated here.

[0276] The computer program product may be implemented in hardware, software or a combination thereof. In an optional embodiment, the computer program product is embodied as a computer storage medium, which may be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK) and the like.

[0277] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the above-described equipment and devices can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0278] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0279] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0280] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0281] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure is described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed in the present disclosure, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be based on the protection scope of the claims.

Claims

1. A SMS AI training model intelligent matching method, characterized in that: include: Obtaining a natural language description, a file name, a file content, and a file format of the data to be processed, and performing semantic segmentation using the natural language description to form structured instructions, and performing feature extraction on the file name, the file content, and the file format to form multi-dimensional feature data; According to the preset model feature portrait library, a feature quadruple dynamic weight matching algorithm is used to process the structured instructions and the multi-dimensional feature data, and a candidate AI training model is output; Based on the candidate AI training model, an intelligent auxiliary decision-making interface is constructed to analyze and display the processing logic flow chart, feature analysis report and historical usage record of the candidate AI training model to generate user feedback results; The user feedback results are used to update the weight parameters in the feature quadruple dynamic weight matching algorithm, and the AI ​​training model in the preset model feature portrait library is re-evaluated and classified according to the updated weight parameters to generate an optimized model feature portrait library.

2. The method according to claim 1, characterized in that The using the natural language description to perform semantic segmentation to form structured instructions includes: By calling a pre-trained scenario-based intent recognition model, the natural language description is divided into multiple semantic blocks; Identify and extract the core action instructions, target objects and data carriers of each semantic block, and obtain preliminary semantic information including the core action instructions, target objects and data carriers; Based on the preliminary semantic information, the core action instructions, the target object and the implementation data carrier are mapped to a preset intention knowledge graph to form the structured instructions.

3. The method according to claim 1, characterized in that The implementation forms multi-dimensional feature data by extracting features from the file name, the file content and the file format, including: Calling a dynamic rule engine to parse the file name, extract time, content, status and format information, and obtain file name features; Performing first-row detection on the file content to identify column header features, using random sampling to determine data patterns, performing format recognition on specific fields, and generating file content features; According to a preset format decoding matrix, the file format is parsed, and a processing method of an encrypted file, a compressed package or unstructured data is determined to form a file format feature; The file name features, file content features and file format features are processed using a feature fusion algorithm to form the multi-dimensional feature data.

4. The method according to claim 1, characterized in that: Through standardized description, the preset model feature profile library is constructed, including: Obtaining functional data of the initial candidate AI training model, and using the scenario tree structure to describe applicable scenarios, and determining the specific application fields and processing objectives of the initial candidate AI training model; Extracting technical parameters including required fields, data formats, and data volume limits based on the technical specifications of the initial candidate AI training model to form input requirement features; Using a standardized capability description language to process the functional characteristics of the initial candidate AI training model, breaking down specific functions into measurable processing units, and generating processing capability characteristics; The specific application field and the processing target, the input requirement characteristics and the processing capability characteristics are integrated to generate a three-dimensional feature file of the candidate AI training model as a preset model feature portrait library.

5. The method according to claim 1, characterized in that The execution process of the feature quadruple dynamic weight matching algorithm includes: Calculate the semantic similarity between the structured instruction and the candidate AI training model function description using the first deep learning model to generate a natural language weight score; Processing key information of the file name features through regular expressions and rule engines, matching it with the target data type of the candidate AI training model, and forming a file name feature weight score; Calculate the matching degree between the file content features and the candidate AI training model input requirements, and output the content feature weight score; Analyze the data format conversion cost corresponding to the file format feature and generate a format adaptation weight score; Using preset weight coefficients, weighted calculation is performed on the natural language weight score, file name feature weight score, content feature weight score and format adaptation weight score to obtain a comprehensive matching degree, which is used as the final scoring result of the candidate AI training model.

6. The method according to claim 1, characterized in that Before executing the feature quadruple dynamic weight matching algorithm, the method further includes: Collect historical matching records, identify high-frequency typical subtask combinations through data mining, build a high-frequency subtask optimization pool, and output task combination patterns; Utilizing the task combination pattern to analyze the execution process of each high-frequency subtask, extracting data dependencies, execution order requirements and resource consumption restrictions, and generating a set of constraint equations; The search space of the candidate AI training models is processed according to the set of constraint equations, and a pruning operation is performed to form an optimized set of candidate AI models.

7. The method according to claim 6, characterized in that The task combination mode is used to analyze the execution process of each high-frequency subtask, extract data dependencies, execution order requirements and resource consumption restrictions, and generate a set of constraint equations, including: Call the first deep learning model to process the model combination pattern in the historical successful cases, and output the model link prediction strategy through pattern recognition and feature learning; Analyze the high-frequency subtasks according to the model link prediction strategy, extract data flow rules and computing resource restriction constraint elements, and form formalized constraint conditions; The constraints are converted into a computational model and a quantifiable set of constraint equations is established to guide subsequent model search and optimization processes.

8. The method according to claim 1, characterized in that Also includes: A feature compression mechanism is adopted in the feature quadruple dynamic weight matching algorithm, including: Obtain the original feature space data, construct a feature pyramid structure through multi-level feature extraction, and output feature combinations at different granularity levels; Based on the feature combination, a top-down iterative matching operation is performed, and the compressed feature data is used to screen the candidate AI training models to generate a preliminary matching result; According to the preliminary matching results, a feature weight adjustment algorithm is constructed to dynamically optimize the feature matching process and output an optimized weight matching algorithm.

9. The method according to claim 8, characterized in that The top-down iterative matching operation is used to select candidate AI training models using compressed feature data, including: Performing preliminary screening through the highest level compressed features of the feature pyramid structure, determining the category range of the candidate AI training models, and outputting model category data; Using the model category data, feature matching is refined in turn at each level of the feature pyramid, and candidate AI training models are determined by screening layer by layer to generate screening results; The screening results are evaluated and verified, and the feature weight combinations and matching strategies whose screening effects reach a preset threshold are stored in the algorithm library to form a reusable feature matching solution.

10. A SMS AI training model intelligent matching system, characterized in that: include: A feature acquisition module, used to acquire a natural language description, a file name, a file content and a file format of the data to be processed, and to perform semantic segmentation using the natural language description to form structured instructions, and to form multi-dimensional feature data by extracting features from the file name, the file content and the file format; A model matching module, used to process the structured instructions and the multi-dimensional feature data using a feature quadruple dynamic weight matching algorithm based on a preset model feature portrait library, and output a candidate AI training model; A decision interface module, which is used to construct an intelligent auxiliary decision interface based on the candidate AI training model, and generate user feedback results by analyzing and displaying the processing logic flow chart, feature analysis report and historical usage record of the candidate AI training model; The optimization and update module is used to use the user feedback results to update the weight parameters in the feature quadruple dynamic weight matching algorithm, re-evaluate and classify the AI ​​training model in the preset model feature portrait library according to the updated weight parameters, and output the optimized model feature portrait library.

Citation Information

Patent Citations

  • Neural network model structure searching method and device, electronic equipment and storage medium

    CN110909877A

  • Data processing method and device, electronic equipment and storage medium

    CN117370373A

  • Neural network model structure determination method and device, equipment, medium and product

    CN117454959A

  • Automatic Detection of Required Tools for a Task Described in Natural Language Content

    US20180157641A1

Cited By

  • Large language model security detection method based on automatic generation of knowledge graph

    CN120180434A

  • A Security Detection Method for Large Language Models Automatically Generated Based on Knowledge Graphs

    CN120180434B

  • Short message template classification and automatic matching method and system based on AI technology

    CN120597863A

  • Dynamic optimization system for AI model training parameters

    CN120633719A

  • Intelligent water meter reading platform and method based on high-precision internet map landmark positioning

    CN121056760A