A supply-demand matching method for personalized services in digital technology
Through natural language big model and word embedding model, deep semantic analysis and vectorization of long technical description texts, combined with cosine similarity and industry label screening, the matching difficulties of cross-industry technical services are solved, and efficient and accurate supply and demand matching effect is achieved.
Patent Information
- Application Number
- CN202510680710.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing supply and demand matching methods have problems with cross-field matching difficulties and insufficient accuracy in cross-industry technical services, especially in complex and multi-dimensional technical service scenarios, which are difficult to meet market demand.
A natural language model is used to understand the long text of technical descriptions in semantics, extract core text and define industry domain labels, and conduct deep semantic vectorization representation through word embedding models. Combining cosine similarity and industry label screening, performance indicator correlation weights are calculated, and the accuracy of matching results is optimized.
It significantly improves the matching accuracy and reliability of cross-industry technical services, ensures accurate matching between the technical supply and demand parties at the performance parameter level, and improves the industry applicability and user experience of the matching results.
Smart Images

Figure CN120196736B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and deep learning matching recommendation, and particularly to a supply-demand matching method for personalized services in digital technology. Background Art
[0002] With the rapid development of digital technology, the accurate matching of technical service capabilities has gradually become an important link in promoting industrial innovation and technology application. In recent years, supply-demand matching methods based on information technology and artificial intelligence have been widely applied in various scenarios.
[0003] CN113724032A discloses a matching method for supply-demand products, which realizes matching by obtaining supply-demand product pairs and calculating product similarity scores and enterprise similarity scores; however, its limitation is that it overemphasizes the product and enterprise attributes of both supply and demand sides, while ignoring the deep semantic features of supply-demand texts and the cross-field matching problems in the industry, making it difficult to achieve efficient and accurate matching of cross-industry technical services.
[0004] CN115660787A discloses an automatic supply-demand precise matching method, which realizes precise supply-demand matching by establishing a demand queue and a supply queue for one-by-one matching; although this method improves the refinement degree of supply-demand matching, in actual operation, it still fails to effectively handle cross-field semantic differences and inconsistencies in industry terms. In addition, this matching method relies on direct keyword matching, which is prone to problems such as one-sidedness or invalidation of retrieval results, such as differences in terms between different fields and ambiguities of the same term in different field contexts, making it difficult to efficiently and accurately identify and match user needs.
[0005] The deficiencies of the above prior arts have led to low utilization efficiency of actual technical service resources and an inability to meet the complex and personalized needs of the current digital technology era, seriously affecting the innovation development and the efficiency and speed of industrial upgrading in the technical service field.
[0006] In summary, the existing supply-demand matching methods have problems of difficult cross-field matching and insufficient accuracy, especially difficult to meet market demands in the face of complex and multi-dimensional technical service scenarios. The present invention provides an efficient and accurate supply-demand matching method for personalized services in digital technology, which fully solves the technical problems of the limitations of traditional keyword matching and insufficient cross-field semantic understanding. Summary of the Invention
[0007] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title, and such simplifications or omissions shall not be used to limit the scope of the present invention.
[0008] In view of the above existing problems, the present invention is proposed.
[0009] To solve the above technical problems, the present invention provides the following technical solutions: for each requirement text in the requirement text set, it is converted into a text vector through a word embedding model, where:
[0010] When the long technical description text x is from the supplier, the text vector is written into the commodity vector database, and the requirement text set is written into the supplier relational database;
[0011] When the long technical description text x is from the demander, the text vector is written into the demand vector database, and the requirement text set is written into the demander relational database;
[0012] For each demand vector in the demand vector database, the cosine similarity is calculated in batches with each commodity vector in the commodity vector database, and the commodity vectors with cosine similarity greater than the similarity threshold enter the first candidate set;
[0013] For each commodity item in the first candidate set, compare its industry field label with the demander label, and if there is no common label, it is excluded, and if there is at least one common label, it is retained;
[0014] Input the demander performance indicators and the retained commodity performance indicators into the natural language large model to obtain the correlation weight between the two;
[0015] Multiply the cosine similarity value by the correlation weight to obtain a similarity score Score, and write it into the matching candidate list in the form of {commodity ID, Score}; if there are duplicate items with the same commodity ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score;
[0016] Sort the matching candidate list according to the final scores from high to low, and push the sorting result to the corresponding client.
[0017] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, obtaining the requirement text set includes:
[0018] Obtain the long technical description text x submitted by the supplier or the demander, call the natural language large model to perform semantic understanding on the long technical description text x, and extract the text Tech that is strongly related to the technical content;
[0019] Input the text Tech into the text classification model, output the probability distribution of each industry field type, and select the top three industry field labels with the highest probability values;
[0020] Call the natural language large model to extract performance index texts containing explicit numerical parameters and performance requirements from the text Tech, and extract a collection of requirement texts from the text Tech according to the technical requirement points.
[0021] As a preferred solution of the supply-demand matching method for digital technology personalized services according to the present invention, the calling of the natural language large model to perform semantic understanding on the long technical description text x and extract the text Tech strongly related to the technical content includes: calling a pre-trained natural language large model, performing semantic understanding on the input long technical description text x through a preset semantic extraction instruction, identifying keywords, key phrases, and core technical paragraphs related to the core technical content in the long text, screening out text fragments strongly related to the technical subject and with consistent semantics, and forming the text Tech.
[0022] As a preferred solution of the supply-demand matching method for digital technology personalized services according to the present invention, the selection of the top three industry field labels with the highest probability values includes:
[0023] Extract features and perform semantic analysis on the text Tech through a deep learning text classification model to generate probability vectors representing each industry field;
[0024] Use the softmax function to normalize the probability scores of each industry field;
[0025] Select the top three labels according to the probability score as the basis for industry field classification.
[0026] As a preferred solution of the supply-demand matching method for digital technology personalized services according to the present invention, the extraction of a collection of requirement texts from the text Tech includes:
[0027] For performance indicators, use the regular expression {parameter name}{comparison operator}{numerical value}{measurement unit} to automatically screen out sentence elements containing numerical constraints;
[0028] For requirement texts, use the natural language large model to execute the technical requirement key point extraction instruction, and divide the text Tech into a collection of requirement sub-clauses with a granularity less than 100 words;
[0029] Return the extraction result in the form of a JSON structure to form a collection of requirement texts.
[0030] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, for each demand text in the demand text set, it is converted into a text vector through a word embedding model, including: using a pre-trained word embedding model to convert each word in the demand text into a vector representation in the semantic space; and then through a text vector aggregation method, the word vectors are concatenated into a unified text vector representing the semantic features of the entire demand text.
[0031] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, calculating the cosine similarity includes: calculating the dot product of the demand text vector and the product text vector, and respectively calculating the norms of their respective vectors, and obtaining the cosine similarity by dividing the dot product by the product of the norms to evaluate the semantic proximity between the two types of vectors.
[0032] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, for each product entry in the first candidate set, comparing its industry field label with the demander label, if there is no common label, it is excluded, and if there is at least one common label, it is retained, including:
[0033] Comparing the intersection of the product entry label set and the demander label set, if the intersection is non-empty, the common label condition is satisfied;
[0034] If the intersection is empty, calculate the Jaccard similarity of the two sets. If the Jaccard similarity ≥ 0.3 and there is an industry synonym mapping relationship, it is regarded as having a common label, otherwise it is regarded as not satisfying the common label condition, and the product entries that do not satisfy the common label condition are removed from the candidate set.
[0035] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, if there are duplicate items with the same product ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score, including:
[0036] A preset bias value β. When a duplicate product ID is detected in the matching candidate list, the bias value β is accumulated into the original score to enhance the sorting priority of the repeated recommendation entries;
[0037] The calculation method of the bias value β is: β = 0.05 × ln(1 + f);
[0038] Where f is the number of repetitions of the same product ID in the matching candidate list, and β is used to enhance the ranking weight of high-frequency candidates while maintaining the stability of the scoring ranking.
[0039] As a preferred solution of the supply-demand matching method for digital technology personalized services described in the present invention, sorting the matching candidate list from high to low according to the final score includes:
[0040] Using the quicksort algorithm to sort the matching candidate list in descending order according to the final score;
[0041] When there are ties in the final scores, first compare the number of common industry tags, and the one with more tags ranks higher;
[0042] If the number of common industry tags is still the same, then sort by the timestamp of commodity information update from the most recent to the oldest;
[0043] Finally, push the top 10 matching results to the corresponding client.
[0044] Advantages of the present invention:
[0045] 1. The present invention uses a natural language large model to perform semantic analysis on the long text of technical descriptions submitted by the supply side or the demand side, extracts the text strongly related to the technical content, realizes the accurate understanding of the text content, solves the limitations of the traditional keyword matching method in semantic understanding, and effectively improves the accuracy and relevance of the preliminary information screening;
[0046] 2. By inputting the extracted text into a text classification model, outputting the probability distribution of each industry field type, and accurately selecting the top three industry field tags with the highest probability values, the term differences across industries are effectively solved, and the accuracy of cross-field technical service supply-demand matching is significantly improved;
[0047] 3. By calling the natural language large model to further extract the performance index text containing explicit numerical parameters and performance requirements, and extracting the demand text set for specific technical requirement points, it ensures that the technical supply and demand sides can be accurately matched at the performance parameter level, and improves the refinement degree of the subsequent matching steps;
[0048] 4. By using the word embedding model to deeply semantically vectorize the demand text set, storing the vectors in the commodity vector database or the demand vector database respectively, it realizes the efficient storage and rapid retrieval of demand and supply information in the semantic space, and ensures the accuracy and efficiency of the subsequent matching calculation;
[0049] 5. By using the method of batch calculating the cosine similarity, it efficiently and accurately realizes the preliminary screening of potential matching items, avoids the redundant interference of irrelevant information, and improves the effectiveness of the preliminary screening;
[0050] 6. Based on the industry field tags for secondary screening, only retaining the technical service items with a high degree of industry field matching greatly improves the industry applicability of the matching results and the user experience;
[0051] 7. Using a natural language large model to calculate the correlation weight between the performance indicators of the demand side and the commodity side accurately reflects the actual correlation degree between performance requirements and supply, making the matching result more valuable for practical applications;
[0052] 8. Calculating the similarity score by comprehensively considering the semantic similarity and the correlation weight of performance indicators fully takes into account the matching degrees in both semantics and performance indicators, greatly improving the accuracy and practicality of the matching result;
[0053] 9. By introducing a preset bias value to handle duplicates, it ensures that the recommended ranking of the same commodity can reasonably reflect its actual matching degree, enhancing the fairness and stability of the matching result;
[0054] 10. Sorting the matching results according to the final score and accurately pushing them to the corresponding client ensures that users can obtain the most relevant technology service recommendations in a timely manner, significantly improving user satisfaction and the efficient allocation efficiency of market resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0056] Figure 1 It is a schematic flowchart of the supply-demand matching method for digital technology personalized services shown in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments.
[0058] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0059] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0060] In view of the above deficiencies of the prior art, the present invention proposes a method for matching the supply and demand of digital technology personalized services based on natural language processing and deep learning. Specifically, the present invention uses a natural language large model to perform semantic understanding and analysis on the long technical text submitted by users, effectively extracts the core text of the technical content, and accurately delimits the industry field tags based on the text classification model, thereby solving the matching obstacles caused by cross-industry term differences;
[0061] In addition, further use the natural language large model to achieve the precise extraction of explicit numerical parameters and performance indicators in the technical text, establish a refined demand text set, and perform deep semantic vectorization representation through the word embedding model; and through efficient cosine similarity calculation and screening rules based on industry tags, the matching accuracy and reliability of cross-industry technical services are greatly improved; at the same time, an associated weight mechanism for performance indicators is proposed, and by comprehensively considering semantic similarity and performance indicator correlation, the accuracy of the matching recommendation result is further optimized.
[0062] According to an embodiment of the present invention, in combination with Figure 1 the flowchart shown, a method for matching the supply and demand of digital technology personalized services specifically includes the following steps:
[0063] S1. Obtain the long technical description text x submitted by the supplier or the demander, and call the natural language large model to perform semantic understanding on the long technical description text x, and extract the text Tech that is strongly related to the technical content. Among them, it should be noted in this step that:
[0064] Call the pre-trained natural language large model, and perform semantic understanding on the input long technical description text x through the preset semantic extraction instruction, identify the keywords, key phrases and core technical paragraphs related to the core technical content in the long text, and screen out the text fragments that are strongly related to the technical subject and have the same semantics to form the text Tech.
[0065] As an example, the specific implementation manner of this step includes:
[0066] Receive the long technical description text x submitted by the supplier or the demander, and this text is .docx 、 .pdf in the form of Markdown or plain text. The system uniformly converts it into UTF-8 encoded plain text and deletes non-technical affiliated information such as headers, footers, and footnotes;
[0067] If the text length exceeds 10,000 characters, it is segmented and cached according to the natural paragraph granularity;
[0068] The server side dynamically generates a semantic extraction instruction for the pre-trained natural language large model (LLM) based on a preset template. An example of this template is as follows:
[0069] "Please extract the paragraphs directly related to the core technical solution from the following text, including key parameters, principles, and implementation methods, and keep the original sentences without rewriting;"
[0070] The semantic extraction instruction contains a keyword weight threshold and a maximum number of returned fragments. For texts detected as bilingual or multilingual, the system will automatically switch the instruction language and threshold settings to adapt to different language environments;
[0071] After receiving the text x and the semantic extraction instruction, the LLM performs word segmentation, sequence encoding, and attention weighting on the text x to extract the Top-K high-weight keywords and key phrases;
[0072] Calculate the cosine similarity between the paragraph vector h p and the keyword centroid vector c. When Sim(h p , c) ≥ the threshold δ, the paragraph is strongly related to the technical subject; if the distance between adjacent strongly related paragraphs does not exceed one natural paragraph, they are merged into the same text fragment to maintain the continuity of the technical context;
[0073] In an optional implementation, the exclusion conditions are: paragraphs containing commercial terms, price information, or non-technical paragraphs that only describe the market prospect are not retained; the retention conditions are: paragraphs containing technical elements such as implementation instructions, system architecture, algorithm flow, key material ratios, or performance parameter thresholds and meeting any of the following rules:
[0074] Contain at least two industry-specific terms;
[0075] Contain at least one performance parameter expression;
[0076] The system encapsulates the filtered text fragments according to a unified JSON structure, and the fields include at least the document ID, original title, fragment index and corresponding text, keyword set, and language identifier;
[0077] The encapsulated text Tech object is written into the "Original-Tech" cache database, and its keyword set is synchronized to the full-text retrieval index for direct invocation by the subsequent industry classification and performance index extraction modules.
[0078] It should be noted that through the implementation of the above steps, the present invention can, while ensuring the integrity of technical semantics, automatically eliminate redundant information, quickly obtain the core text Tech that is strongly related to the technical subject and has consistent semantics, provide high-quality data input for subsequent industry label determination, performance index extraction, and vectorization modeling, thereby significantly improving the overall accuracy of supply-demand matching and the robustness of the system.
[0079] S2. Input the text Tech into the text classification model, output the probability distribution of each industry field type, and select the top three industry field labels with the highest probability values. It should be noted in this step that:
[0080] Perform sub-word granularity word segmentation on the text Tech obtained in step S1, remove stop words and retain technical proper nouns, and generate an ordered token sequence of length L;
[0081] Call the multi-layer Transformer-Encoder network to generate d-dimensional context-related vectors for each token mapping , and combine all to form an embedding matrix;
[0082] Calculate the attention weights on the embedding matrix according to the industry prior dictionary to obtain the industry-weighted sentence vector v. At the same time, introduce the global mean pooling vector g, and cascade the two and input them into the differentiable kernel function layer to achieve cross-domain feature fusion;
[0083] Use the deep learning text classification model to calculate the normalized probabilities of each industry field :
[0084] ;
[0085] Among them, is the normalized probability of industry i, and the value range is (0, 1). is the d-dimensional context vector of the j-th sub-word token in the embedding space. is the differentiable attention kernel function. is the multi-scale gating mapping function. is the learnable parameter of industry i, I is the total number of system-preset industry categories, and L is the length of the text Tech sub-word sequence.
[0086] According to sort in descending order, and select the top three industry field labels with the highest probability values as the industry classification labels of the text Tech for label consistency comparison in step S6.
[0087] It should be noted that when and the output of is non-negative and , the exponential term in the numerator is always positive, so , when indicates that the text Tech highly confidently belongs to a single industry; indicates that there are cross-industry cross features and need to be comprehensively judged in combination with the Top-3 labels; indicates that the text expression is ambiguous or the industry span is too large and it is necessary to go back to steps S1~S4 to optimize keyword extraction.
[0088] S3. Call the natural language large model to extract the performance index text containing explicit numerical parameters and performance requirements from the text Tech, and extract the requirement text set from the text Tech according to the technical requirement points. It should be noted in this step that:
[0089] Divide the text Tech into sentence element sequences according to natural paragraphs, with the length limit of each sentence element not exceeding 150 words. Generate a unique index for each paragraph of sentence elements and write it into the temporary corpus buffer for regular matching and semantic reasoning calls;
[0090] For the sentence element sequence, apply regular expressions in the form of 、 、 、 where represents the technical parameter name, represents the comparison operator (such as ≥, ≤, >, <, =), represents the decimal value, represents the measurement unit;
[0091] Mark the successfully matched sentence elements as performance index type and write them into the performance index list Perf;
[0092] Batch input the sentence elements not marked as performance index type into the LLM and execute the preset requirement point extraction instruction. The instruction requires the LLM to extract the content directly related to technical problems, functional objectives or implementation conditions as requirement sub-clauses not exceeding 100 words without changing the original sentence word order. The set of extracted sub-clauses is denoted as Req;
[0093] Perform semantic similarity deduplication on the sentence elements in Perf and Req respectively, with the threshold set to 0.85. If the adjacent requirement sub-clauses are separated by ≤ 1 sentence and semantically coherent, they are merged into a complete requirement expression to avoid splitting the technical context;
[0094] Concatenate the results after denoising and merging into a unified JSON structure to obtain the requirement text set:
[0095] ;
[0096] where is the requirement text set, is the k-th requirement entry in the set, is the entry category identifier (performance or requirement), is the original index of the corresponding sentence element of the entry in the text Tech, is the original sentence element text content without rewriting. N is the total number of entries in the requirement text set, and write this structure into the requirement extraction result cache library for step S4 to call.
[0097] S4. For each requirement text in the requirement text set, convert it into a text vector through a word embedding model. It should be noted in this step that:
[0098] For the requirement text set Perform sub-word granularity tokenization one by one, remove stop words and retain technical proper nouns to obtain a word sequence of length M ;
[0099] Call the pre-trained word embedding model Map each word to a d-dimensional context-related vector , and cache it in the temporary tensor buffer;
[0100] According to the word embedding aggregation formula that fits the document semantics, perform integral pooling and energy normalization processing on in the high-dimensional semantic space to generate a unified text vector V representing the semantic features of the entire requirement text:
[0101]
[0102] where V is the text vector obtained after aggregation, M is the number of valid word tokens in the requirement text, is the semantic attention coefficient of the i-th word, is the hyperbolic tangent function, is the position-sensitive scaling constant, is the word 's d-dimensional embedding vector, is the continuous integral variable, is the second polylogarithmic integral function, is the global attenuation coefficient positive number; is the Euclidean norm, is the smoothing constant positive number to prevent the denominator from being zero;
[0103] When is close to 1, it indicates that the internal semantics of the text is highly consistent and the technical key points are concentrated; when approaches the lower limit, it indicates that the text semantics is discrete, and it is necessary to go back to step S3 to re-extract the requirement clauses;
[0104] When the long technical description text x is from the supplier, write the text vector into the commodity vector database and write the requirement text set into the supplier relational database;
[0105] When the long technical description text x is from the demander, write the text vector into the demand vector database and write the requirement text set into the demander relational database.
[0106] It should be noted that this step not only preserves the word-level context dependency information, but also realizes robust filtering of long text noise by using integral pooling and multi-logarithmic energy suppression. At the same time, a sub-library writing architecture is adopted for the text vectors of the supply side and the demand side, providing a unified and highly confident vector representation for step S5, thus significantly improving the accuracy of supply-demand matching in cross-industry scenarios.
[0107] S5. For each demand vector in the demand vector database, batch calculate the cosine similarity with each commodity vector in the commodity vector database, and the commodity vectors with cosine similarity greater than the similarity threshold enter the first candidate set. Among them, it should be noted in this step that:
[0108] Read the demand text vectors one by one from the demand vector database while constructing a batch candidate cache based on the ball tree index in the commodity vector database ;
[0109] For any pair Call the following formula to calculate the improved cosine similarity The mathematical expression formula is as follows:
[0110]
[0111] Among them, is the cosine similarity, Q is a positive integer of the text vector dimension, is the i-th dimensional semantic refraction coefficient, the measurable real number of the i-th component of the demand vector, is the measurable real number of the i-th component of the commodity vector, is the Gaussian error function, is a positive real number of the position attenuation factor, is the continuous integral variable, is the exponential integral function, is a positive real number of the difference attenuation coefficient, is the modulus length exponential adjustment factor, is the modulus length smoothing constant;
[0112] Compare with the threshold If then write into the first candidate set and record the corresponding demand ID. For example, when it indicates a very high semantic fit between the demand and the commodity;
[0113] If , such as , it is regarded as low correlation and filtered;
[0114] After completing a round of batch calculations, write the first candidate set in the form of triples into the intermediate result buffer for step S6 to call.
[0115] S6. For each commodity entry in the first candidate set, compare its industry field label with the demander's label. If there is no common label, remove it; if there is at least one common label, keep it. Here, it should be noted in this step that:
[0116] Compare the intersection of the commodity entry label set and the demander's label set. If the intersection is not empty, the common label condition is met;
[0117] If the intersection is empty, calculate the Jaccard similarity of the two sets. If the Jaccard similarity ≥ 0.3 and there is an industry synonym mapping relationship, it is considered that there is a common label; otherwise, it is considered that the common label condition is not met, and the commodity entries that do not meet the common label condition are removed from the candidate set.
[0118] S7. Input the demander's performance indicators and the retained commodity performance indicators into the natural language large model to obtain the correlation weight between the two. Here, it should be noted in this step that:
[0119] Normalize the same dimension of the commodity performance indicators retained in step S6 and their corresponding demander's performance indicators, convert different measurement units into a unified expression, and map the range description (such as ≥ 90%) to a closed interval;
[0120] For each group of demand - commodity indicators, perform hash matching with the indicator name as the key to form a <demand value, commodity value> paired structure. If the commodity lacks the corresponding indicator, fill in a null value at this position and add a default mark;
[0121] Generate the LLM input text according to the preset prompt template, which includes the indicator name, two - end values, comparison operators, and importance weight prompts to ensure that the corpus fully presents the demand - commodity differences and business background;
[0122] Call the pre - trained natural language large model, assign the evaluation task instruction in the system role, inject the context generated in the previous step in the user role, and the model outputs a continuous real number , where indicates that the commodity indicator meets the demand, indicates that the commodity indicator does not meet the demand;
[0123] Map to through a piece - wise linear function, and introduce a confidence reduction coefficient to suppress the high variance of the model and obtain the final correlation weight ;
[0124] Map Form a triple <demand value, product value> with the demand ID and the product ID Write it into the weight cache table for step S8 to call and multiply with the cosine similarity to generate the comprehensive similarity score.
[0125] It should be noted that, on the premise of ensuring measurement consistency, this step makes full use of the semantic inference ability of the natural language large model for complex performance descriptions, obtains objective, quantifiable and controllable range of correlation weights, provides a high-confidence index-level compensation for the subsequent matching score calculation, and thus significantly improves the accuracy and robustness of cross-industry supply and demand matching.
[0126] S8. Multiply the cosine similarity value by the correlation weight to obtain the similarity score Score, and write it into the matching candidate list in the form of {product ID, Score}. Among them, what needs to be explained in this step is:
[0127] (1) Read one by one from the intermediate result buffer of steps S5 to S7 quadruple, where is the cosine similarity,[[]] is the index correlation weight;
[0128] (2) Perform [0,1] linear stretching on and respectively, map the extreme values to the closed interval endpoints to ensure the stability of subsequent integration;
[0129] (3) Call the following formula to generate the comprehensive similarity score S:
[0130]
[0131] where S is the comprehensive similarity score, is the cosine similarity,[[]] is the index correlation weight, is the integral decay constant positive real number, is the integral independent variable, D is the summation upper limit, n is the summation index natural number, is the dilogarithm integral function, is the denominator smoothing constant positive real number;
[0132] When it means that the product technical performance highly matches the demand and can be directly recommended;
[0133] When it means medium matching and needs to be secondarily screened in combination with the label weight;
[0134] When it is regarded as low matching and a bias value is added in step S9;
[0135] (4) According to the triple Write the matching candidate list for steps S9 - S10 to call;
[0136] (5) When the calculated S does not satisfy the interval [0, 1], automatically backtrack to step (2) for recalibration And Limit it to 1.
[0137] It should be noted that the implementation of this step, while maintaining the dual constraints of semantic similarity and performance weight, introduces an integral decay and high - order polylogarithmic compensation mechanism, achieving two - way suppression of local overfitting and high - dimensional noise, providing a smoother and more discriminative comprehensive score for subsequent sorting output, thus significantly improving the accuracy and robustness of cross - industry supply - demand matching.
[0138] S9. If there are duplicate items with the same product ID in the matching candidate list, then a preset bias value is superimposed on the original similarity score Score. Among them, it should be noted in this step that:
[0139] The preset bias value β. When a duplicate product ID is detected in the matching candidate list, the bias value β is accumulated into the original score to enhance the sorting priority of duplicate recommendation entries;
[0140] The calculation method of the bias value β is:
[0141] Among them, f is the number of repetitions of the same product ID in the matching candidate list, and β is used to enhance the ranking weight of high - frequency candidates while maintaining the stability of the scoring ranking.
[0142] It should be noted that if there are no duplicate items with the same product ID in the matching candidate list, the similarity score Score calculated in step S8 is used as the final score for step S10 to perform sorting processing.
[0143] S10. Sort the matching candidate list in descending order according to the final score and push the sorting result to the corresponding client. Among them, it should be noted in this step that:
[0144] Use the quicksort algorithm to sort the matching candidate list in descending order according to the final score;
[0145] When there are ties in the final scores, first compare the number of common industry tags, and the one with more tags ranks higher;
[0146] If the number of common industry tags is still the same, then sort by the product information update timestamp from the most recent to the oldest;
[0147] Finally, push the top 10 matching results to the corresponding client.
[0148] It should also be noted that by introducing a multi-level collaborative mechanism of semantic understanding - industry determination - performance index linkage in the supply-demand matching process, the present invention realizes the full-stack parsing and constraint of technical texts from the macro industry context to the micro numerical parameters, producing unexpected technical effects that exceed the existing keyword retrieval or simple similarity ranking methods;
[0149] On the one hand, the embodiment of the present invention uses a natural language large model to perform deep semantic extraction on the long text Tech and combines it with probabilistic industry classification, which can actively discover cross-domain synonyms, antonyms, and implicit associations, enabling the accurate recall of technology supply-demand pairs that were previously difficult to capture between different industrial vertical fields; on the other hand, the dual numerical constraints formed by the regular screening of performance indicators and the weight inference of the large model endow the matching score with a quantitatively controllable "soft-hard combination" ability, neither overly excluding solutions with potentially adjustable indicators nor effectively filtering candidates with significantly substandard performance, significantly improving the application availability and actual implementation success rate of the matching results.
[0150] Preferably, in the comprehensive similarity calculation stage, the present invention couples the cosine similarity with a multi-order logarithmic integral compensation term and superimposes a logarithmic bias dynamic correction based on the occurrence frequency, which not only solves the problem of score jitter caused by high-dimensional vector noise amplification but also forms a gradient ranking effect within the candidate set where multi-dimensional satisfaction is strengthened and single-dimensional shortcomings are weakened. Combined with the secondary filtering of the industry label intersection and the Jaccard threshold, the recall rate is accurately maintained.
[0151] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A supply-demand matching method for personalized services in digital technology, characterized in that, Including: For each requirement text in the requirement text set, convert it into a text vector through a word embedding model, where: When the long text x of the technical description is from the supplier, write the text vector into the commodity vector database, and write the set of demand texts into the supplier relational database; When the long technical description text x is from the requester, write the text vector into the requirement vector database, and write the set of requirement texts into the relational database of the requester; For each requirement vector in the requirement vector database, batch calculate the cosine similarity with each commodity vector in the commodity vector database, and the commodity vectors with a cosine similarity greater than the similarity threshold enter the first candidate set; For each commodity entry in the first candidate set, compare its industry field label with the requester label, and remove it if there is no common label, and retain it if there is at least one common label; Input the requester performance indicators and the retained commodity performance indicator pairs into a large natural language model to obtain the correlation weight between them; where: Perform dimension normalization on the retained commodity performance indicators and their corresponding requester performance indicators, and convert different measurement units into a unified expression; For each pair of the performance indicators of the demand side and the retained product performance indicators, perform a hash match with the indicator name as the key to form a paired structure. If the product lacks the corresponding indicator, fill in a null value at that position and add a default marker; Generate the LLM input text according to a preset prompt template, and this template includes the indicator name, two-end value, comparison operator, and importance weight; Call the pre-trained large natural language model, assign an evaluation task instruction to the system role, inject the context generated in the previous step into the user role, and the model outputs continuous real numbers , where indicates that the product indicators meet the requirements, indicates that the product indicators do not meet the requirements; Be mapped to through a piecewise linear function, and introduce a confidence reduction coefficient to suppress the high variance of the model and obtain the correlation weight ; Multiply the cosine similarity value by the correlation weight to obtain the similarity score Score, and write it into the matching candidate list in the form of {commodity ID, Score}; If there are duplicate items with the same commodity ID in the matching candidate list, superimpose a preset bias value on the original similarity score Score; Sort the matching candidate list according to the final score from high to low, and push the sorting result to the corresponding client.
2. The supply-demand matching method for digital technology personalized services according to claim 1, characterized in that, Obtain the requirement text set, including: Obtain the long technical description text x submitted by the supplier or the requester, call the large natural language model to perform semantic understanding on the long technical description text x, and extract the text Tech that is strongly related to the technical content; Input the text Tech into a text classification model, output the probability distribution of each industry field type, and select the top three industry field labels with the highest probability values; Call the large natural language model to extract the performance indicator text containing explicit numerical parameters and performance requirements from the text Tech, and extract the requirement text set from the text Tech according to the technical requirement points.
3. The supply-demand matching method for digital technology personalized services according to claim 2, characterized in that, The call to the large natural language model to perform semantic understanding on the long technical description text x and extract the text Tech that is strongly related to the technical content includes: calling a pre-trained large natural language model, performing semantic understanding on the input long technical description text x through a preset semantic extraction instruction, identifying keywords, key phrases, and core technical paragraphs related to the core technical content in the long text, screening out text fragments that are strongly related to the technical subject and have consistent semantics, and forming the text Tech.
4. The supply-demand matching method for digital technology personalized services according to claim 2, characterized in that, The selection of the top three industry field labels with the highest probability values includes: Extract features and perform semantic analysis on the text Tech through a deep learning text classification model to generate a probability vector representing each industry field; Normalize the probability scores of each industry field using the softmax function; Select the top three labels according to the probability score as the basis for industry field classification.
5. The supply-demand matching method for digital technology personalized services according to claim 2, wherein Extract the requirement text set from the text Tech, including: For performance indicators, use the regular expression {parameter name}{comparison operator}{numerical value}{measurement unit} to automatically screen out sentence elements containing numerical constraints; For the requirement text, execute the technical requirement key point extraction instruction using a natural language large model, and split the text Tech into a set of requirement sub-clauses with a granularity less than 100 words; Return the extraction result in the form of a JSON structure to form a collection of requirement texts.
6. The supply-demand matching method for digital technology personalized services according to claim 1 or 5, characterized in that For each requirement text in the collection of requirement texts, convert it into a text vector through a word embedding model, including: using a pre-trained word embedding model to convert each word in the requirement text into a vector representation in the semantic space; then, through a text vector aggregation method, splice the word vectors into a unified text vector representing the semantic features of the entire requirement text.
7. The supply-demand matching method for digital technology personalized services according to claim 1, wherein Calculate the cosine similarity, including: calculating the dot product of the requirement text vector and the product text vector, and calculating the norm lengths of their respective vectors separately, and obtaining the cosine similarity by dividing the dot product by the product of the norm lengths to evaluate the semantic proximity between the two types of vectors.
8. The supply-demand matching method for digital technology personalized services according to claim 7, characterized in that, For each product item in the first candidate set, compare its industry domain label with the requester's label. If there is no common label, it is removed; if there is at least one common label, it is retained, including: Compare the intersection of the product item label set and the requester's label set. If the intersection is non-empty, the common label condition is satisfied; If the intersection is empty, calculate the Jaccard similarity of the two sets. If the Jaccard similarity ≥ 0.3 and there is an industry synonym mapping relationship, it is considered that there is a common label; otherwise, it is considered that the common label condition is not satisfied, and the product item that does not satisfy the common label condition is removed from the candidate set.
9. The supply-demand matching method for digital technology personalized services according to claim 1, characterized in that If there are duplicate items with the same product ID in the matching candidate list, a preset bias value is superimposed on the original similarity score Score, including: A preset bias value β. When a duplicate product ID is detected in the matching candidate list, the bias value β is accumulated into the original score to enhance the sorting priority of the duplicate recommended items; The calculation method of the bias value β is: β = 0.05 × ln(1 + f) where f is the number of repetitions of the same product ID in the matching candidate list, and β is used to improve the ranking weight of high-frequency candidates while maintaining the stability of the score ranking.
10. The supply-demand matching method for digital technology personalized services according to claim 9, characterized in that Sort the matching candidate list according to the final score from high to low, including: Use the quicksort algorithm to sort the matching candidate list in descending order according to the final score; When there are ties in the final scores, first compare the number of common industry labels, and the one with more labels ranks higher; If the number of common industry labels is still the same, then sort by the product information update timestamp from the most recent to the oldest; Finally, push the top 10 matching results to the corresponding client.
Citation Information
Patent Citations
Supply and demand product matching method and device, electronic equipment and storage medium
CN113724032A
Query request response method and device
CN115114424A
Content recommendation method, system and equipment and storage medium
CN116521985A