Supply and demand matching method for digital science and technology personalized service

Semantic understanding and industry label demarcation through natural language big models, combined with word embedding model and cosine similarity calculation, the problems of insufficient semantic understanding and insufficient accuracy in cross-domain matching are solved, and efficient and accurate matching of supply and demand of technical services are achieved.

CN120196736AActive Publication Date: 2025-06-24JIANGSU PRODUCTIVITY PROMOTION CENT

Patent Information

Application Number
CN202510680710.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing supply and demand matching methods have problems of insufficient semantic understanding and insufficient accuracy when matching across fields, and it is difficult to meet the complex and personalized technical service needs.

Method used

A natural language model is used to semantically understand long texts of technical descriptions, extract texts that are strongly related to technical content, and accurately define industry domain labels through text classification models. Deep semantic vectorization is performed through the word embedding model, cosine similarity is calculated, and the weights are combined with industry labels and performance indicators to generate a comprehensive similarity score.

Benefits of technology

It significantly improves the accuracy and reliability of the supply and demand matching of cross-industry technical services, ensures that both the technical supply and demand parties can accurately match at the semantic and performance parameters level, and improves user satisfaction and the effective allocation efficiency of market resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196736A_ABST
    Figure CN120196736A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing and deep learning matching recommendation, and discloses a supply and demand matching method for digital science and technology personalized services, which comprises the following steps: carrying out semantic understanding and analysis on a submitted technical long text by adopting a natural language large model, effectively extracting a core text of technical contents, and carrying out semantic analysis on the core text; and based on the text classification model, accurately delimiting industry field labels, thereby solving matching obstacles caused by cross-industry term differences. Precise extraction of clear numerical parameters and performance indexes in technical texts is realized by using a natural language large model, a refined demand text set is established, and deep semantic vectorization representation is performed through a word embedding model; through cosine similarity calculation and a screening rule based on industry labels, the matching accuracy and reliability of cross-industry technical services are greatly improved; based on an association weight mechanism of performance indexes, semantic similarity and performance index association degree are comprehensively considered, and the accuracy of matching recommendation results is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing and deep learning matching recommendation, and particularly to a supply-demand matching method for personalized services in digital technology. Background Art

[0002] With the rapid development of digital technology, the precise matching of technical service capabilities has gradually become an important link in promoting industrial innovation and technological application. In recent years, supply-demand matching methods based on information technology and artificial intelligence have been widely applied in various scenarios.

[0003] CN113724032A discloses a matching method for supply-demand products, which realizes matching by obtaining supply-demand product pairs and calculating product similarity scores and enterprise similarity scores; however, its limitation lies in overemphasizing the product and enterprise attributes of both supply and demand sides while ignoring the deep semantic features of supply-demand texts and cross-domain matching problems in the industry, making it difficult to achieve efficient and precise matching for cross-industry technical services.

[0004] CN115660787A discloses an automatic supply-demand precise matching method, which realizes precise supply-demand matching by establishing a demand queue and a supply queue for one-by-one matching; although this method improves the refinement degree of supply-demand matching, it still fails to effectively handle cross-domain semantic differences and inconsistencies in industry terms in actual operation. In addition, this matching method relies on direct keyword matching, which is prone to problems such as one-sidedness or invalidation of retrieval results, such as differences in terms between different fields and ambiguities of the same term in different field contexts, making it difficult to efficiently and accurately identify and match user needs.

[0005] The above deficiencies of the prior art have led to low utilization efficiency of actual technical service resources and inability to meet the complex and personalized needs of the current digital technology era, seriously affecting the innovation development and industrial upgrading efficiency and speed in the technical service field.

[0006] In summary, the existing supply-demand matching methods have problems of difficult cross-domain matching and insufficient accuracy, especially difficult to meet market demands in the face of complex and multi-dimensional technical service scenarios. The present invention provides an efficient and precise supply-demand matching method for personalized services in digital technology, fully solving the technical problems of the limitations of traditional keyword matching and insufficient cross-domain semantic understanding. Summary of the Invention

[0007] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the specification of this application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions cannot be used to limit the scope of the present invention.

[0008] In view of the above existing problems, the present invention is proposed.

[0009] To solve the above technical problems, the present invention provides the following technical solutions: For each requirement text in the requirement text set, it is converted into a text vector through a word embedding model, where: When the long technical description text x is from the supplier, the text vector is written into the commodity vector database, and the requirement text set is written into the supplier relational database; When the long technical description text x is from the demander, the text vector is written into the demand vector database, and the requirement text set is written into the demander relational database; For each demand vector in the demand vector database, the cosine similarity is calculated in batches with each commodity vector in the commodity vector database, and the commodity vectors with a cosine similarity greater than the similarity threshold enter the first candidate set; For each commodity item in the first candidate set, compare its industry field label with the demander label. If there is no common label, it is excluded. If there is at least one common label, it is retained; Input the demander performance index and the retained commodity performance index pair into the large natural language model to obtain the correlation weight between the two; Multiply the cosine similarity value by the correlation weight to obtain the similarity score Score, and write it into the matching candidate list in the form of {commodity ID, Score}; if there are duplicate items with the same commodity ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score; Sort the matching candidate list according to the final scores from high to low, and push the sorting result to the corresponding client.

[0010] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, obtaining the requirement text set includes: Obtain the long technical description text x submitted by the supplier or the demander, call the large natural language model to perform semantic understanding on the long technical description text x, and extract the text Tech that is strongly related to the technical content; Input the text Tech into the text classification model, output the probability distribution of each industry field type, and select the top three industry field labels with the highest probability values; Call the large natural language model to extract the performance index text containing explicit numerical parameters and performance requirements from the text Tech, and extract the requirement text set from the text Tech according to the technical requirement points.

[0011] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, the step of invoking a large natural language model to perform semantic understanding on the long technical description text x and extracting the text Tech strongly related to the technical content includes: invoking a pre-trained large natural language model, performing semantic understanding on the input long technical description text x through a preset semantic extraction instruction, identifying keywords, key phrases, and core technical paragraphs related to the core technical content in the long text, screening out text fragments strongly related to the technical subject and with consistent semantics, and forming the text Tech.

[0012] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, the step of selecting the top three industry field tags with the highest probability values includes: Performing feature extraction and semantic analysis on the text Tech through a deep learning text classification model to generate a probability vector representing each industry field; Normalizing the probability scores of each industry field using the softmax function; Selecting the top three tags according to the probability scores as the basis for industry field classification.

[0013] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, the step of extracting a demand text set from the text Tech includes: For performance indicators, using a regular expression {parameter name}{comparison operator}{numerical value}{measurement unit} to automatically screen out sentence elements containing numerical constraints; For demand texts, using a large natural language model to execute a technical demand key point extraction instruction to split the text Tech into a set of demand sub-clauses with a granularity less than 100 words; Returning the extraction result in the form of a JSON structure to form a demand text set.

[0014] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, converting each demand text in the demand text set into a text vector through a word embedding model includes: using a pre-trained word embedding model to convert each word in the demand text into a vector representation in the semantic space; and then through a text vector aggregation method, splicing the word vectors into a unified text vector representing the semantic features of the entire demand text.

[0015] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, calculating the cosine similarity includes: calculating the dot product of the demand text vector and the commodity text vector, and respectively calculating the norms of their respective vectors, and obtaining the cosine similarity by dividing the dot product by the product of the norms to evaluate the semantic proximity between the two types of vectors.

[0016] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, for each commodity entry in the first candidate set, compare its industry field label with the demander label. If there is no common label, it is excluded; if there is at least one common label, it is retained, including: Compare the intersection of the commodity entry label set and the demander label set. If the intersection is non-empty, the common label condition is satisfied; If the intersection is empty, calculate the Jaccard similarity of the two sets. If the Jaccard similarity ≥ 0.3 and there is an industry synonym mapping relationship, it is considered that there is a common label; otherwise, it is considered that the common label condition is not satisfied, and the commodity entries that do not meet the common label condition are removed from the candidate set.

[0017] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, if there are duplicate items with the same commodity ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score, including: Preset a bias value β. When a duplicate commodity ID is detected in the matching candidate list, the bias value β is accumulated into the original score to enhance the sorting priority of the duplicate recommendation entries; The calculation method of the bias value β is: β = 0.05 × ln(1 + f); where f is the number of repetitions of the same commodity ID in the matching candidate list, and β is used to improve the ranking weight of high-frequency candidates while maintaining the stability of the score ranking.

[0018] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, sort the matching candidate list from high to low according to the final score, including: Use the quicksort algorithm to sort the matching candidate list in descending order according to the final score; When there are tied final scores, first compare the number of common industry labels, and the one with more labels ranks higher; If the number of common industry labels is still the same, then sort by the commodity information update timestamp from recent to far; Finally, push the top 10 matching results to the corresponding client.

[0019] Advantages of the present invention: 1. The present invention uses a natural language large model to perform semantic analysis on the long text of technical descriptions submitted by the supplier or demander, extracts the text strongly related to the technical content, realizes the accurate understanding of the text content, solves the limitations of the traditional keyword matching method in semantic understanding, and effectively improves the accuracy and relevance of the preliminary information screening; 2. By inputting the extracted text into a text classification model, the probability distribution of each industry field type is output, and the top three industry field labels with the highest probability values are accurately selected, effectively solving the term differences across industries and significantly improving the accuracy of the supply-demand matching of cross-field technical services; 3. By calling a large natural language model to further extract performance index texts containing explicit numerical parameters and performance requirements, and extracting a collection of requirement texts for specific technical requirement points, it ensures that both the technical supply and demand sides can be accurately matched at the performance parameter level, improving the refinement degree of subsequent matching steps; 4. Through the word embedding model, the collection of requirement texts is deeply semantically vectorized, and the vectors are respectively stored in the commodity vector database or the demand vector database, realizing the efficient storage and fast retrieval of demand and supply information in the semantic space, and ensuring the accuracy and efficiency of subsequent matching calculations; 5. By adopting the method of batch calculating the cosine similarity, the preliminary screening of potential matching items is efficiently and accurately realized, avoiding the redundant interference of irrelevant information and improving the effectiveness of the preliminary screening; 6. Based on the industry field labels, secondary screening is carried out, and only the technical service entries with a high degree of industry field matching are retained, greatly improving the industry applicability of the matching results and the user experience; 7. By using a large natural language model to calculate the correlation weight between the performance indicators of the demand side and the commodity side, it accurately reflects the actual correlation degree between the performance demand and the supply, making the matching results more valuable for practical applications; 8. By comprehensively calculating the similarity score by combining the semantic similarity and the performance indicator correlation weight, it fully considers the matching degrees of both semantics and performance indicators, greatly improving the accuracy and practicality of the matching results; 9. By introducing a preset bias value to handle duplicates, it ensures that the recommended ranking of the same commodity can reasonably reflect its actual matching degree, improving the fairness and stability of the matching results; 10. According to the final score, the matching results are sorted and accurately pushed to the corresponding client, ensuring that users can obtain the most relevant technical service recommendations in a timely manner, significantly improving user satisfaction and the efficient allocation efficiency of market resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them: Figure 1 It is a schematic flowchart of the supply-demand matching method for digital technology personalized services shown in the present invention. Detailed implementation manners

[0021] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.

[0022] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0024] In view of the above-mentioned deficiencies of the prior art, the present invention proposes a method for matching supply and demand of digital technology personalized services based on natural language processing and deep learning. Specifically, the present invention uses a natural language large model to perform semantic understanding and analysis on the long technical text submitted by users, effectively extracts the core text of the technical content, and accurately delimits industry domain tags based on the text classification model, thereby solving the matching obstacles caused by cross-industry term differences; In addition, further use the natural language large model to achieve accurate extraction of explicit numerical parameters and performance indicators in the technical text, establish a refined demand text set, and perform deep semantic vectorization representation through the word embedding model; and through efficient cosine similarity calculation and screening rules based on industry tags, the matching accuracy and reliability of cross-industry technical services are greatly improved; at the same time, an associated weight mechanism for performance indicators is proposed, and by comprehensively considering semantic similarity and performance indicator correlation, the accuracy of the matching recommendation results is further optimized.

[0025] According to the embodiments of the present invention, in combination with Figure 1 the flowchart shown, a method for matching supply and demand of digital technology personalized services specifically includes the following steps: S1. Obtain the long technical description text x submitted by the supplier or the demander, and call the natural language large model to perform semantic understanding on the long technical description text x, and extract the text Tech that is strongly related to the technical content. Among them, it should be noted in this step that: Call the pre-trained natural language large model, perform semantic understanding on the input long technical description text x through the preset semantic extraction instruction, identify the keywords, key phrases, and core technical paragraphs related to the core technical content in the long text, and screen out the text fragments that are strongly related to the technical subject and have the same semantics to form the text Tech.

[0026] As an example, the specific implementation of this step includes: Receiving the long text x of the technical description submitted by the supplier or the demander, which is in the form of .docx , .pdf , Markdown or plain text. The system uniformly converts it into UTF-8 encoded plain text and deletes non-technical ancillary information such as headers, footers, and footnotes; If the text length exceeds 10,000 characters, it is segmented and cached by natural paragraph granularity; The server side dynamically generates semantic extraction instructions for the pre-trained large natural language model (LLM) based on a preset template. An example of this template is as follows: "Please extract the paragraphs directly related to the core technical solution from the following text, including key parameters, principles, and implementation methods, and keep the original sentences without rewriting;" The semantic extraction instructions contain keyword weight thresholds and the maximum number of returned segments. For texts detected as bilingual or multilingual, the system will automatically switch the instruction language and threshold settings to adapt to different language environments; After receiving the text x and the semantic extraction instructions, the LLM performs word segmentation, sequence encoding, and attention weighting on the text x, and extracts the Top-K high-weight keywords and key phrases; Calculate the cosine similarity between the paragraph vector h p and the keyword centroid vector c. When Sim(h p , c) ≥ the threshold δ, this paragraph is strongly related to the technical subject. If the distance between adjacent strongly related paragraphs does not exceed one natural paragraph, they are merged into the same text segment to maintain the continuity of the technical context; In an alternative implementation, the exclusion conditions are: non-technical paragraphs containing commercial terms, price information, or only describing the market prospect are not retained; the retention conditions are: paragraphs containing technical elements such as implementation instructions, system architecture, algorithm flow, key material ratios, or performance parameter thresholds and meeting any of the following rules: Containing at least two industry-specific terms; Containing at least one performance parameter expression; The system encapsulates the filtered text segments in a unified JSON structure, and the fields include at least the document ID, original title, segment index and corresponding text, keyword set, and language identifier; The encapsulated text Tech object is written into the "Original-Tech" cache database, and its keyword set is synchronized to the full-text retrieval index for direct invocation by the subsequent industry classification and performance index extraction modules.

[0027] It should be noted that through the implementation of the above steps, the present invention can automatically eliminate redundant information while ensuring the integrity of technical semantics, quickly obtain the core text Tech that is strongly relevant to the technical subject and has consistent semantics, provide high-quality data input for subsequent industry label determination, performance index extraction, and vectorization modeling, thereby significantly improving the overall accuracy of supply-demand matching and the robustness of the system.

[0028] S2. Input the text Tech into the text classification model, output the probability distribution of each industry field type, and select the top three industry field labels with the highest probability values. It should be noted in this step that: Perform sub-word granularity word segmentation on the text Tech obtained in step S1, remove stop words and retain technical proper nouns, and generate an ordered token sequence of length L; Call the multi-layer Transformer-Encoder network to generate d-dimensional context-related vectors for each token mapping , and all to form an embedding matrix; Calculate the attention weights on the embedding matrix according to the industry prior dictionary to obtain the industry-weighted sentence vector v. At the same time, introduce the global mean pooling vector g, and cascade the two and input them into the differentiable kernel function layer to achieve cross-domain feature fusion; Use the deep learning text classification model to calculate the normalized probabilities of each industry field : ; Among them, is the normalized probability of industry i, with a value range of (0, 1), is the d-dimensional context vector of the j-th sub-word token in the embedding space, is the differentiable attention kernel function, is the multi-scale gating mapping function, is the learnable parameter of industry i, I is the total number of system preset industry categories, and L is the length of the sub-word sequence of text Tech.

[0029] According to sort in descending order, and select the top three industry field labels with the highest probability values as the industry classification labels of text Tech for label consistency comparison in step S6.

[0030] It should be noted that when and the output of is non-negative and , the exponential term in the numerator is always positive, so , when it means that text Tech highly confidently belongs to a single industry; Indicates that there are cross - industry cross - features and need to be comprehensively judged in combination with the Top - 3 tags; It indicates that the text expression is ambiguous or the industry span is too large, and it is necessary to backtrack to steps S1 - S4 to optimize keyword extraction.

[0031] S3. Call the large natural language model to extract the performance index text containing clear numerical parameters and performance requirements from the text Tech, and extract the requirement text set from the text Tech according to the technical requirement points. Among them, it should be noted in this step that: Divide the text Tech into sentence element sequences according to natural paragraphs. The length of each sentence element is limited to no more than 150 words. Generate a unique index for each paragraph of sentence elements and write it into the temporary corpus buffer for regular matching and semantic reasoning calls; For the sentence element sequence, apply regular expressions in the form of 、 、 、 where, represents the technical parameter name, represents the comparison operator (such as ≥, ≤, >, <, =), represents the decimal value, represents the measurement unit; Mark the successfully matched sentence elements as performance index type and write them into the performance index list Perf; Batch - input the sentence elements not marked as performance index type into the LLM, and execute the preset requirement key point extraction instruction. The instruction requires the LLM to extract the content directly related to technical problems, functional objectives or implementation conditions as requirement sub - clauses not exceeding 100 words without changing the original sentence word order. The set of extracted sub - clauses is denoted as Req; Perform semantic similarity deduplication on the sentence elements in Perf and Req respectively, with a threshold of 0.85. If the adjacent requirement sub - clauses are separated by ≤ 1 sentence and semantically coherent, they are merged into a complete requirement expression to avoid splitting the technical context; Concatenate the results after denoising and merging into a unified JSON structure to obtain the requirement text set: ; where, is the requirement text set, is the k - th requirement entry in the set, is the entry category identifier (performance or requirement), is the original index of the corresponding sentence element of the entry in the text Tech, is the original sentence element text content without rewriting, N is the total number of entries in the requirement text set, and write this structure into the requirement extraction result cache library for step S4 to call.

[0032] S4. For each requirement text in the requirement text set, convert it into a text vector through a word embedding model. It should be noted that in this step: For the requirement text set Perform sub-word granularity word segmentation item by item, remove stop words and retain technical proper nouns to obtain a word sequence of length M ; Call the pre-trained word embedding model Map each word to a d-dimensional context-related vector , and cache it in the temporary tensor buffer; According to the word embedding aggregation formula that fits the document semantics, perform integral pooling and energy normalization processing on in the high-dimensional semantic space to generate a unified text vector V representing the semantic features of the entire requirement text:

[0033] where V is the text vector obtained after aggregation, M is the number of valid word tokens in the requirement text, is the semantic attention coefficient of the i-th word, is the hyperbolic tangent function, is the position-sensitive scaling constant, is the word 's d-dimensional embedding vector, is the continuous integral variable, is the second polylogarithmic integral function, is the global decay coefficient positive number; is the Euclidean norm, is the smoothing constant positive number to prevent the denominator from being zero; When is close to 1, it indicates that the internal semantics of the text is highly consistent and the technical key points are concentrated; when approaches the lower limit, it indicates that the text semantics is discrete, and it is necessary to go back to step S3 to re-extract the requirement clauses; When the long technical description text x originates from the supplier, write the text vector into the product vector database and write the requirement text set into the supplier relational database; When the long technical description text x originates from the requester, write the text vector into the requirement vector database and write the requirement text set into the requester relational database.

[0034] It should be noted that this step not only retains the word-level context dependency information, but also uses integral pooling and multi-logarithmic energy suppression to achieve robust filtering of long text noise. At the same time, a separate library writing architecture is adopted for the text vectors of the supplier and demander, which provides a unified and high-confidence vector representation for step S5, thereby significantly improving the accuracy of supply and demand matching in cross-industry scenarios.

[0035] S5. For each demand vector in the demand vector database, the cosine similarity is calculated in batches with each product vector in the product vector database, and the product vectors whose cosine similarity is greater than the similarity threshold enter the first candidate set. Among them, what needs to be explained in this step is: Read the requirement text vectors one by one in the requirement vector database At the same time, a batch candidate cache based on ball tree index is constructed in the product vector database ; For any pair Call the following formula to calculate the improved cosine similarity , its mathematical expression formula is as follows:

[0036] in, is the cosine similarity, Q is a positive integer of the text vector dimension, is the semantic refraction coefficient of the i-th dimension, The i-th component of the demand vector is a measurable real number, is a measurable real number of the ith component of the commodity vector, is the Gaussian error function, is a positive real number for the position attenuation factor, is a continuous integral variable, is the exponential integral function, is a positive real number of the difference attenuation coefficient, is the modulus length index adjustment factor, is the modulus smoothing constant; Will With threshold Compare, if Then Write the first candidate set and record the corresponding requirement ID. Indicates that the demand and product semantics are highly consistent; like ,like , it is considered as low correlation and filtered; After completing a round of batch calculation, the first candidate set is The triplet form is written into the intermediate result buffer for calling by step S6.

[0037] S6. For each product entry in the first candidate set, compare its industry field label with the label of the demander. If there is no common label, it is removed; if there is at least one common label, it is retained. It should be noted in this step that: Compare the intersection of the product entry label set and the demander label set. If the intersection is non-empty, the common label condition is met; If the intersection is empty, calculate the Jaccard similarity of the two sets. If the Jaccard similarity ≥ 0.3 and there is an industry synonym mapping relationship, it is considered that there is a common label; otherwise, it is considered that the common label condition is not met, and the product entries that do not meet the common label condition are removed from the candidate set.

[0038] S7. Input the demander performance indicators and the retained product performance indicators into the natural language large model to obtain the correlation weight between the two. It should be noted in this step that: Perform dimension normalization on the product performance indicators retained in step S6 and their corresponding demander performance indicators, convert different measurement units into a unified expression, and map range descriptions (such as ≥ 90%) to closed intervals; For each group of demand-product indicators, perform hash matching with the indicator name as the key to form a <demand value, product value> paired structure. If the product is missing the corresponding indicator, fill in a null value at this position and add a default mark; Generate the LLM input text according to the preset prompt template. This template includes the indicator name, two-end value, comparison operator, and importance weight prompt to ensure that the corpus completely presents the demand-product difference and business background; Call the pre-trained natural language large model, assign an evaluation task instruction in the system role, and inject the context generated in the previous step in the user role. The model outputs a continuous real number , where indicates that the product indicator meets the demand, indicates that the product indicator does not meet the demand; Map to through a piecewise linear function, and introduce a confidence reduction coefficient to suppress the high variance of the model and obtain the final correlation weight ; Combine with the demand ID and product ID to form a triple <demand value, product value> and write it into the weight cache table for step S8 to call and generate a comprehensive similarity score by multiplying with the cosine similarity.

[0039] It should be noted that, on the premise of ensuring measurement consistency, this step makes full use of the semantic inference ability of the large natural language model for complex performance descriptions to obtain the correlation weights within the objective, quantifiable and controllable range, providing a high-confidence index-level compensation for the subsequent matching score calculation, thus significantly improving the accuracy and robustness of cross-industry supply-demand matching.

[0040] S8. Multiply the cosine similarity value by the correlation weight to obtain the similarity score Score, and write it into the matching candidate list in the form of {commodity ID, Score}. It should be noted in this step that: (1) Read the quadruples one by one from the intermediate result buffer of steps S5 to S7, where is the cosine similarity, and is the index correlation weight; (2) Perform [0,1] linear stretching on and respectively to map the extreme values to the closed interval endpoints to ensure the stability of subsequent integration; (3) Call the following formula to generate the comprehensive similarity score S:

[0041] where S is the comprehensive similarity score, is the cosine similarity, is the index correlation weight, is the integral decay constant positive real number, is the integral independent variable, D is the summation upper limit, n is the summation index natural number, is the dilogarithm integral function of the second order, is the denominator smoothing constant positive real number; When it means that the technical performance of the commodity highly matches the demand and can be directly recommended; When it means that the medium match requires secondary screening in combination with the label weight; When it is regarded as a low match and a bias value is added in step S9; (4) Write into the matching candidate list according to the triple for steps S9 to S10 to call; (5) When the calculated S does not satisfy the [0,1] interval, automatically backtrack to step (2) to recalibrate and and limit it to 1.

[0042] It should be noted that in the implementation of this step, while maintaining the dual constraints of semantic similarity and performance weight, an integral decay and high-order polylogarithmic compensation mechanism is introduced to achieve two-way suppression of local overfitting and high-dimensional noise, providing a smoother and more discriminative comprehensive score for subsequent sorting output, thereby significantly improving the accuracy and robustness of cross-industry supply-demand matching.

[0043] S9. If there are duplicate items with the same product ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score. Here, it should be noted in this step that: The preset bias value β. When a duplicate product ID is detected in the matching candidate list, the bias value β is accumulated into the original score to enhance the sorting priority of the repeated recommendation entries; The calculation method of the bias value β is:

[0044] where f is the number of repetitions of the same product ID in the matching candidate list, and β is used to enhance the ranking weight of high-frequency candidates while maintaining the stability of the score ranking.

[0045] It should be noted that if there are no duplicate items with the same product ID in the matching candidate list, the similarity score Score calculated in step S8 is used as the final score for the sorting process in step S10.

[0046] S10. Sort the matching candidate list in descending order according to the final score and push the sorting result to the corresponding client. Here, it should be noted in this step that: The quicksort algorithm is used to sort the matching candidate list in descending order according to the final score; When there are ties in the final scores, the number of common industry labels is compared first, and the one with more labels ranks higher; If the number of common industry labels is still the same, then sort by the product information update timestamp from the most recent to the oldest; Finally, the top 10 matching results are pushed to the corresponding client.

[0047] It should also be noted that the present invention realizes a full-stack parsing and constraint of technical texts from the macro industry context to the micro numerical parameters by introducing a multi-level collaborative mechanism of semantic understanding - industry determination - performance index linkage in the supply-demand matching process, producing unexpected technical effects that exceed the existing keyword retrieval or simple similarity sorting methods; On the one hand, the embodiment of the present invention utilizes a large natural language model to perform deep semantic extraction on the long text Tech and combines it with probabilistic industry classification, which can actively discover cross-domain synonyms, heteronyms and implicit associations, so that the technology supply and demand pairs that were previously difficult to capture between different industry verticals can also be accurately recalled; on the other hand, the dual numerical constraints formed by regular screening of performance indicators and weight inference of large models give quantitative and controllable "soft and hard" capabilities in matching scores, which neither excessively excludes solutions with potential adjustable indicators, nor can it effectively filter out candidates whose performance is obviously substandard, significantly improving the application availability and actual implementation success rate of matching results.

[0048] Preferably, the present invention couples cosine similarity with multi-order logarithmic integral compensation terms in the comprehensive similarity obtaining stage, and superimposes logarithmic bias dynamic correction based on the frequency of occurrence, which not only solves the problem of score jitter caused by high-dimensional vector noise amplification, but also forms a gradient sorting effect within the candidate set, which strengthens multi-dimensional satisfaction and weakens single-dimensional shortcomings. Combined with the secondary filtering of industry label intersection and Jaccard threshold, the recall rate is accurately maintained.

[0049] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A supply-demand matching method for personalized services in digital technology, characterized in that, Including: For each requirement text in the requirement text set, convert it into a text vector through a word embedding model, where: When the long technical description text x is from the supplier, write the text vector into the commodity vector database, and write the set of demand texts into the supplier relational database; When the long text x of the technical description is from the requester, write the text vector into the requirement vector database, and write the set of requirement texts into the relational database of the requester; For each requirement vector in the requirement vector database, batch calculate the cosine similarity with each commodity vector in the commodity vector database, and the commodity vectors with a cosine similarity greater than the similarity threshold enter the first candidate set; For each commodity entry in the first candidate set, compare its industry field label with the requester label. If there is no common label, it is excluded; if there is at least one common label, it is retained; Input the requester performance indicators and the retained commodity performance indicator pairs into a large natural language model to obtain the correlation weight between them; Multiply the cosine similarity value by the correlation weight to obtain a similarity score Score, and write it into the matching candidate list in the form of {commodity ID, Score}; If there are duplicate items with the same commodity ID in the matching candidate list, add a preset offset value to the original similarity score Score; Sort the matching candidate list according to the final scores from high to low, and push the sorting result to the corresponding client.

2. The supply-demand matching method for digital technology personalized services according to claim 1, wherein Obtain the requirement text set, including: Obtain the long technical description text x submitted by the supplier or the requester, call a large natural language model to perform semantic understanding on the long technical description text x, and extract the text Tech that is strongly related to the technical content; Input the text Tech into a text classification model, output the probability distribution of each industry field type, and select the top three industry field labels with the highest probability values; Call the large natural language model, extract the performance indicator text containing explicit numerical parameters and performance requirements from the text Tech, and extract the requirement text set from the text Tech according to the technical requirement points.

3. The supply-demand matching method for digital technology personalized services according to claim 2, wherein The step of calling a large natural language model to perform semantic understanding on the long technical description text x and extract the text Tech that is strongly related to the technical content includes: calling a pre-trained large natural language model, performing semantic understanding on the input long technical description text x through a preset semantic extraction instruction, identifying keywords, key phrases, and core technical paragraphs related to the core technical content in the long text, screening out text fragments that are strongly related to the technical subject and have consistent semantics, and forming the text Tech.

4. The supply-demand matching method for digital technology personalized services according to claim 2, characterized in that The step of selecting the top three industry field labels with the highest probability values includes: Performing feature extraction and semantic analysis on the text Tech through a deep learning text classification model to generate a probability vector representing each industry field; Normalizing the probability scores of each industry field using the softmax function; Selecting the top three labels according to the probability scores as the basis for industry field classification.

5. The supply-demand matching method for digital technology personalized services according to claim 2, wherein Extracting the requirement text set from the text Tech includes: For performance indicators, use the regular expression {parameter name}{comparison operator}{numerical value}{measurement unit} to automatically screen out sentence elements containing numerical constraints; For requirement texts, use a large natural language model to execute a technical requirement key point extraction instruction, and split the text Tech into a set of requirement sub-clauses with a granularity less than 100 words; Return the extraction result in the form of a JSON structure to form a requirement text set.

6. The supply-demand matching method for digital technology personalized services according to claim 1 or 5, characterized in that For each requirement text in the set of requirement texts, convert it into a text vector through a word embedding model, including: using a pre-trained word embedding model to convert each word in the requirement text into a vector representation in the semantic space; and then through a text vector aggregation method, concatenate the word vectors into a unified text vector representing the semantic features of the entire requirement text.

7. The supply-demand matching method for digital technology personalized services according to claim 1, characterized in that Calculate the cosine similarity, including: calculating the dot product of the requirement text vector and the product text vector, and respectively calculating the norms of their respective vectors, and obtaining the cosine similarity by dividing the dot product by the product of the norms, so as to evaluate the semantic proximity between the two types of vectors.

8. The supply-demand matching method for digital technology personalized services according to claim 7, characterized in that For each product entry in the first candidate set, compare its industry field label with the requester's label. If there is no common label, it is removed. If there is at least one common label, it is retained, including: Compare the intersection of the product entry label set and the requester's label set. If the intersection is non-empty, the common label condition is satisfied; If the intersection is empty, calculate the Jaccard similarity of the two sets. If the Jaccard similarity ≥ 0.3 and there is an industry synonym mapping relationship, it is regarded as having a common label, otherwise it is regarded as not satisfying the common label condition, and the product entries that do not satisfy the common label condition are removed from the candidate set.

9. The supply-demand matching method for digital technology personalized services according to claim 1, wherein If there are duplicate entries with the same product ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score, including: Preset the bias value β. When a duplicate product ID is detected in the matching candidate list, the bias value β is accumulated into the original score to enhance the sorting priority of the duplicate recommendation entries; The calculation method of the bias value β is: β = 0.05 × ln(1 + f); where f is the number of repetitions of the same product ID in the matching candidate list, and β is used to improve the ranking weight of high-frequency candidates while maintaining the stability of the score ranking.

10. The supply-demand matching method for digital technology personalized services according to claim 9, characterized in that, Sort the matching candidate list from high to low according to the final score, including: Use the quicksort algorithm to sort the matching candidate list in descending order according to the final score; When there are ties in the final scores, first compare the number of common industry labels, and the one with more labels ranks higher; If the number of common industry labels is still the same, then sort by the product information update timestamp from the most recent to the oldest; Finally, push the top 10 matching results to the corresponding client.

Citation Information

Patent Citations

  • Supply and demand product matching method and device, electronic equipment and storage medium

    CN113724032A

  • Query request response method and device

    CN115114424A

  • Content recommendation method, system and equipment and storage medium

    CN116521985A

  • Human-in -loop artificial intelligence classification

    WO2024249322A1

Cited By

  • Engineering quantity list intelligent matching method

    CN120634593A

  • Intelligent matching method for bill of quantities

    CN120634593B

  • Welfare scheme selection device based on large model and supply chain

    CN120851825A

  • Urban governance data search method and device based on large model technology

    CN121188267A

  • Supplier analysis method, electronic device and computer program product

    CN121210634A