Target information screening method and system
By leveraging machine learning algorithms and large-scale model technology, combined with multiple evaluation models and pre-feature screening modules, efficient and low-cost target information screening is achieved. This solves the problems of screening efficiency and accuracy of traditional methods in big data environments, and improves the accuracy and cost-effectiveness of specific target information screening.
Patent Information
- Application Number
- CN202511074730.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional data filtering methods are inefficient when dealing with complex and variable data, and are difficult to adapt to dynamic changes in data. Furthermore, in the context of big data, the filtering cost is high and the accuracy is low, making it difficult to meet the needs of filtering specific target information.
By employing machine learning algorithms and large model technology, target filtering parameters are obtained, and multiple evaluation models and pre-feature filtering modules are used to generate query instructions and input them into the target large model to achieve the filtering of target information.
It significantly reduces the cost of filtering target information, improves the accuracy of filtering and the efficiency of information acquisition, especially in specific target information scenarios, it improves the accuracy and cost-effectiveness of filtering results.
Smart Images

Figure CN120974007A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and specifically to a method and system for filtering target information. Background Technology
[0002] The explosive growth of data volume has made it crucial to sift through massive amounts of data to extract target information. Traditional data filtering and information acquisition methods often rely on manually set rules or simple statistical methods, which are inefficient when dealing with complex and variable data and are difficult to adapt to dynamic changes in the data.
[0003] The emergence of big data technology and AI big data models offers solutions to this problem. By collecting, storing, processing, and analyzing massive amounts of data, and leveraging the data understanding capabilities of AI big data models, the dimensionality and accuracy of problem analysis are effectively improved, and the efficiency of information acquisition is enhanced.
[0004] However, as the amount of data increases significantly, the difficulty and cost of filtering out the information of interest also increase, while the accuracy of the filtering results decreases. For specific target information filtering scenarios, such as filtering key technical information or key project information, low-cost, high-efficiency, high-accuracy, and highly objective filtering tools are needed to meet the filtering and decision-making requirements. In this case, traditional solutions show obvious shortcomings. Summary of the Invention
[0005] To address the aforementioned limitations, this invention proposes a target information filtering method and system that utilizes machine learning algorithms and large-scale modeling techniques to obtain target information filtering results. This system can reduce the cost of target information filtering and improve the filtering accuracy and information acquisition cost in specific target information filtering scenarios.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A method for filtering target information, the method comprising the following steps:
[0008] Step 1: Obtain target filtering parameters; these parameters are used to constrain the target range, number of results, and filtering precision of the target filtering.
[0009] Step 2: Based on the target filtering parameters, obtain the first clue data and the second clue data;
[0010] Step 3: Input the first clue data into the first evaluation model and the second evaluation model respectively to obtain the first evaluation result and the second evaluation result; input the second clue data into the third evaluation model to obtain the third evaluation result;
[0011] Step 4: Calculate the fourth evaluation result based on the first and second evaluation results;
[0012] Step 5: Input the third and fourth evaluation results and target screening parameters into the pre-feature screening module to obtain the third clue data;
[0013] Step 6: Generate several query commands from the third-party data, input them into the target model, and summarize the target information filtering results.
[0014] Compared with the prior art, the present invention has the following advantages:
[0015] (1) Preliminary screening is carried out by using pre-assessment and screening methods, which significantly reduces the cost of screening target information using large models;
[0016] (2) Improve the accuracy of target information screening results by evaluating and calculating comprehensive value and overall value;
[0017] (3) Using unstructured text data as an evaluation factor and multimodal data collaborative processing improves the accuracy of screening.
[0018] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the steps of a target information filtering method provided in an embodiment of the present invention.
[0020] Figure 2 This is a structural diagram of a target information filtering system provided in an embodiment of the present invention. Detailed Implementation
[0021] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. To further understand the present invention, the present invention will be further described in detail below with reference to the preferred embodiments.
[0022] One aspect of the present invention is a target information filtering method, referring to Figure 1 The method includes the following steps:
[0023] Step 1: Obtain target filtering parameters; these parameters are used to constrain the target range, number of results, and filtering precision of the target filtering.
[0024] Step 2: Based on the target filtering parameters, obtain the first clue data and the second clue data;
[0025] Step 3: Input the first clue data into the first evaluation model and the second evaluation model respectively to obtain the first evaluation result and the second evaluation result; input the second clue data into the third evaluation model to obtain the third evaluation result;
[0026] Step 4: Calculate the fourth evaluation result based on the first and second evaluation results;
[0027] Step 5: Input the third and fourth evaluation results and target screening parameters into the pre-feature screening module to obtain the third clue data;
[0028] Step 6: Generate several query commands from the third-party data, input them into the target model, and summarize the target information filtering results.
[0029] Another aspect of the present invention is a target information filtering system, referring to Figure 2 The system consists of a parameter input module, a data collection module, a pre-evaluation module, a pre-feature screening module, an instruction generation module, a large model module, and a result generation module.
[0030] The parameter input module is used to obtain target filtering parameters and determine the compliance of the target filtering parameters;
[0031] The data collection module is used to generate a data query request based on the target filtering parameters and to obtain first clue data and second clue data from a remote server.
[0032] The pre-evaluation module is used to obtain the first evaluation result, the second evaluation result, and the third evaluation result by means of the first evaluation model, the second evaluation model, and the third evaluation model, and to calculate the fourth evaluation result;
[0033] The pre-feature screening module is used to obtain third clue data based on the third evaluation result, the fourth evaluation result, and the target screening parameters;
[0034] The instruction generation module is used to generate several query instructions based on the third clue data;
[0035] The large model module is used to generate corresponding query results by calling the target large model that can be queried online according to the query command;
[0036] The result generation module is used to combine the results generated by the large model module into target information filtering results.
[0037] As one embodiment, the target filtering parameters include target object, target range code, time range, target country code, maximum number of results, and filtering accuracy.
[0038] The target filtering parameters are obtained from the parameter input module.
[0039] As one embodiment, step 2 specifically includes the following steps:
[0040] Step 21: Convert the target range code, time range, and target country code in the target filtering parameters into corresponding request parameters according to the preset parameter conversion rules, and form a data request instruction;
[0041] Step 22: The data collection module sends the data request instruction to the remote server and receives the data in response from the remote server;
[0042] Step 23: Perform data cleaning and text preprocessing on the data responded by the remote server to obtain the first clue data and the second clue data.
[0043] The first clue data contains several structured feature data, and the feature data includes at least one target object description field that can characterize the target object.
[0044] The second clue data is obtained by preprocessing the text data in response from the remote server. The preprocessing includes text cleaning, word segmentation, and entity recognition. Each piece of preprocessed data contains a target object description field.
[0045] The first clue data and the second clue data are associated through the target object description field.
[0046] As one embodiment, if the first clue data is patent feature data, it should at least consist of the following data fields: patent number, applicant, inventor, country of publication, country of applicant, patent type, patent status, patent application date, patent publication date, number of patent families, number of transfers, number of licenses, and number of pledges.
[0047] The second clue data is the patent text data after text preprocessing.
[0048] As one embodiment, the first evaluation model and the second evaluation model are models trained based on machine learning algorithms;
[0049] The first evaluation model is used to predict importance features; the second evaluation model is used to predict value features.
[0050] The first evaluation model is obtained by training an importance prediction training set using a machine learning algorithm; the importance prediction training set consists of several feature data and corresponding importance scores.
[0051] The importance score is a numerical value in the range [0,1]. The smaller the value, the lower the importance, and vice versa.
[0052] The second evaluation model is obtained by training a value prediction training set using a machine learning algorithm; the value prediction training set consists of several feature data and corresponding value score vectors.
[0053] The elements in the value score vector are all values in the range [0,1]. Each element represents a value score obtained based on one or more corresponding indicators. The smaller the value, the lower the value, and vice versa.
[0054] The machine learning algorithm is implemented using a neural network algorithm. The specific training method is a mature existing technology, which can be successfully implemented by those skilled in the art based on the description of the foregoing embodiments, and will not be described in detail here.
[0055] It is understood that the first assessment result consists of several importance score prediction values; the second assessment result consists of several value score prediction value vectors, in the form of: [x1,x2,x3,x4,…,x n ].
[0056] As one embodiment, the third evaluation model is used to predict the value features of text data. It is trained based on the Transformer algorithm and consists of a training set composed of text data and corresponding value score vectors. The method of training the model using the Transformer algorithm is a mature existing technology and will not be described in detail here.
[0057] The third evaluation result consists of several text value vectors, each representing a score for a text value dimension. The number of elements contained in the text value vector is equal to the number of data entries in the second clue data.
[0058] As one embodiment, step 4 specifically includes the following steps:
[0059] Step 41: Dynamically update the weight data of value elements;
[0060] Step 42: Calculate the fourth evaluation result based on the updated value element weights and the first and second evaluation results; the fourth evaluation result consists of several feature data and corresponding comprehensive value scores.
[0061] Furthermore, in step 41, the method for dynamically updating the value element weight data includes:
[0062] Step 411: Initialize element weight data, specifically including:
[0063] The weights of each element in the element weight data are reset according to a uniform distribution, with each element having a weight of 1. Where n is the number of elements in the value score prediction vector;
[0064] Step 412: Calculate the value prediction error output by the second evaluation model based on the historical evaluation data output by the second evaluation model.
[0065] Step 413: Dynamically adjust the weight data of the value elements using the Q-learning algorithm.
[0066] In step 412, the specific calculation method for the value prediction error is as follows:
[0067]
[0068] Among them, e j Let x be the mean absolute error of the j-th element. ij It is the predicted value of the i-th value score in the historical assessment data. is the actual labeled value of the i-th value score, and m is the number of entries in the first clue data.
[0069] Step 413 specifically includes the following steps:
[0070] Step 4131: Map the value prediction error obtained in step 412 into discrete states (including three states: low error, medium error, and high error) according to a preset mapping rule;
[0071] Step 4132: Set the Q value of each discrete state-action pair to 0 to obtain an initial Q value table; it is understood that the actions include increasing weight, decreasing weight, and maintaining weight;
[0072] Step 4133: Calculate the instant reward. The calculation method is as follows:
[0073]
[0074] Where eprev is the error of the previous round in the historical evaluation data;
[0075] Step 4134: Independently update the Q-value for each value weight element unit. Iterate until the Q-value change converges or the preset maximum number of iterations is reached. Update the Q-value as follows during each update:
[0076] Q(s,a)+η[R+γ·max a′ Q(s′,a′)-Q(s,a)]
[0077] Where Q(s,a) represents the Q-value of the current discrete state-action pair, η is the preset learning rate, γ is the preset discount factor, and max a' Q(s',a') represents the maximum Q value of the next discrete state;
[0078] Step 4135: After the iterative update is completed, update the corresponding value weight elements based on the updated Q-value; the updated action is selected as the action with the largest Q-value in the updated discrete state; the updated value weight elements are:
[0079]
[0080] Where, α j (t+1) For the updated value weight element values, α j (t+1) The value of the element before the update is Δα, which is the preset update step size. The value of sign(a*) is determined according to the following rules: when the update action is to increase the weight, sign(a*) is 1; when the update action is to decrease the weight, sign(a*) is -1; when the update action is to maintain the weight, sign(a*) is 0.
[0081] It should be noted that after all the value weight elements have been updated, the values of each value weight element need to be normalized.
[0082] In step 42, the comprehensive value score is calculated as follows:
[0083]
[0084] Among them, v i For the i-th comprehensive value score, w i Let α be the predicted value of the i-th importance score. j Let x be the value element weight of the j-th element in the value score prediction vector. ij Let j be the value of the j-th element in the vector of predicted values for the i-th value score.
[0085] The element weights are adjusted and calculated in real time based on the data distribution to improve model adaptability. The Q-learning algorithm provides self-learning capabilities, reduces manual intervention, and improves subsequent screening accuracy.
[0086] As one embodiment, step 5 specifically includes:
[0087] Step 51: Match and group the third and fourth evaluation results according to their target object description fields;
[0088] Step 52: Calculate the overall value score for each group;
[0089] Step 53: Sort each group in reverse order according to the overall value score, truncate the results according to the maximum number of results in the target screening parameters, and extract all data from the target object description field to obtain the third clue data.
[0090] Furthermore, the overall value score is calculated as follows:
[0091]
[0092] Among them, u p For the overall value score of group p, w pq v is the predicted importance score for the q-th element in the p-th group. pq For the q-th comprehensive value score in the p-th group, TextScore p γ represents the average value score of the text data within the p-th group, and γ is the preset global text weight coefficient.
[0093] The average value score for text data is calculated as follows:
[0094]
[0095] Where, r p Let β be the number of data points contained in the p-th group. k Y represents the weight of the k-th text value dimension. pk This represents the total score for the k-th text value dimension in the p-th group.
[0096] As one embodiment, the text value dimension weights are obtained in the following way:
[0097] (1) Select m samples from the historical data of the third evaluation data, and calculate the initial weight of the text value dimension using the entropy weight method. The calculation method is as follows:
[0098]
[0099] Where, β k (0) Let y be the initial weight of the k-th text value dimension. i k Score the k-th text value dimension for the i-th data sample;
[0100] (2) Continuously optimize the weighting of the text value dimension through a sliding window feedback mechanism:
[0101] Whenever the amount of text data processed reaches the preset number of text entries, the prediction error of the weight of each text value dimension is calculated;
[0102] When the prediction error exceeds the preset weight prediction error threshold, the corresponding text value dimension weight is increased according to the preset ratio.
[0103] The weights of all text value dimensions are normalized.
[0104] As one embodiment, the query instruction is generated in step 6 in the following way:
[0105] The third-party clue data is populated into a template generated according to preset instructions to obtain the query instructions.
[0106] As one embodiment, in step 6, after generating the query instruction, the query instruction is input into the large model module, and the large model module sends the query instruction to the remote AI large model server in the form of an API request and receives the response.
[0107] As one embodiment, the method described in this invention can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device.
[0108] The method described in this invention can be implemented in the form of a software program, which can be executed by a processor to achieve the steps or functions described above. Similarly, the software program (including associated data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices.
[0109] In addition, some steps or functions of the method described in this invention can be implemented in hardware, for example, as a circuit that works with a processor to perform the various steps or functions.
[0110] Furthermore, a portion of the methods described in this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. The program instructions invoking the methods described in this invention can be stored in a fixed or removable recording medium, and / or transmitted via a data stream in a broadcast or other signal carrying medium, and / or stored in the working memory of a computer device operating according to the program instructions.
[0111] As one embodiment, the present invention also provides an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing plurality of embodiments.
[0112] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0113] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0114] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0115] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for filtering target information, characterized in that, The method includes the following steps: Step 1: Obtain target filtering parameters; Step 2: Based on the target filtering parameters, obtain the first clue data and the second clue data; Step 3: Input the first clue data into the first evaluation model and the second evaluation model respectively to obtain the first evaluation result and the second evaluation result; input the second clue data into the third evaluation model to obtain the third evaluation result; Step 4: Calculate the fourth evaluation result based on the first and second evaluation results; Step 5: Input the third and fourth evaluation results and target screening parameters into the pre-feature screening module to obtain the third clue data; Step 6: Generate several query commands from the third-party data, input them into the target model, and summarize the target information filtering results.
2. The method according to claim 1, characterized in that, The target filtering parameters are used to constrain the target range, number of results, and filtering accuracy of the target filtering. The target filtering parameters include target object, target range code, time range, target country code, maximum number of results, and filtering accuracy.
3. The method according to claim 1, characterized in that, Step 21: Convert the target range code, time range, and target country code in the target filtering parameters into corresponding request parameters according to the preset parameter conversion rules, and form a data request instruction; Step 22: The data collection module sends the data request instruction to the remote server and receives the data in response from the remote server; Step 23: Perform data cleaning and text preprocessing on the data responded by the remote server to obtain the first clue data and the second clue data.
4. The method according to claim 1, characterized in that, The first evaluation model and the second evaluation model are models trained based on machine learning algorithms; The first evaluation model is used to predict importance characteristics; the second evaluation model is used to predict value characteristics. The first evaluation model is trained using a machine learning algorithm on an importance prediction training set; the importance prediction training set consists of several feature data and corresponding importance scores; The second evaluation model is obtained by training a value prediction training set using a machine learning algorithm; the value prediction training set consists of several feature data and corresponding value score vectors. The machine learning algorithm is implemented using a neural network algorithm.
5. The method according to claim 1, characterized in that, Step 4 specifically includes the following steps: Step 41: Dynamically update the weight data of value elements; Step 42: Calculate the fourth evaluation result based on the updated value element weights and the first and second evaluation results; The fourth evaluation result consists of several characteristic data and corresponding comprehensive value scores.
6. The method according to claim 5, characterized in that, The method for dynamically updating the weight data of the value elements includes: Step 411: Initialize element weight data; Step 412: Calculate the value prediction error output by the second evaluation model based on the historical evaluation data output by the second evaluation model. Step 413: Dynamically adjust the weight data of the value elements using the Q-learning algorithm.
7. The method according to claim 5, characterized in that, The comprehensive value score is calculated as follows: Among them, v i For the i-th comprehensive value score, w i Let α be the predicted value of the i-th importance score. j Let x be the value element weight of the j-th element in the value score prediction vector. ij Let j be the value of the j-th element in the vector of predicted values for the i-th value score.
8. The method according to claim 1, characterized in that, Step 5 specifically includes: Step 51: Match and group the third evaluation result and the fourth evaluation result according to the target object description field of their feature data; Step 52: Calculate the overall value score for each group; Step 53: Sort each group in reverse order according to the overall value score, truncate the results according to the maximum number of results in the target screening parameters, and extract all data in the target object description field to obtain the second clue data and the third clue data. The overall value score is calculated as follows: Among them, u p For the overall value score of group p, w pq v is the predicted importance score for the q-th element in the p-th group. pq For the q-th comprehensive value score in the p-th group, TextScore p γ is the average value score of the text data in the p-th group, and γ is the preset global text weight coefficient. The average value score for text data is calculated as follows: Where, r p Let β be the number of data points contained in the p-th group. k Y represents the weight of the k-th text value dimension. pk This represents the total score for the k-th text value dimension in the p-th group.
9. The method according to claim 8, characterized in that, The weights of the text value dimension are obtained in the following way: (1) Select m samples from the historical data of the third evaluation data, and calculate the initial weight of the text value dimension using the entropy weight method. The calculation method is as follows: Where, β k (0) Let y be the initial weight of the k-th text value dimension. i k Score the k-th text value dimension for the i-th data sample; (2) Continuously optimize the weighting of the text value dimension through a sliding window feedback mechanism: Whenever the amount of text data processed reaches the preset number of text entries, the prediction error of the weight of each text value dimension is calculated; When the prediction error exceeds the preset weight prediction error threshold, the corresponding text value dimension weight is increased according to the preset ratio. The weights of all text value dimensions are normalized.
10. A target information filtering system, characterized in that, The system consists of the following modules: The parameter input module is used to obtain the target filtering parameters and determine the compliance of the target filtering parameters; The data collection module is used to generate data query requests based on target filtering parameters and to obtain first-line and second-line data from a remote server. The pre-evaluation module is used to obtain the first evaluation result, the second evaluation result, and the third evaluation result by means of the first evaluation model, the second evaluation model, and the third evaluation model, and to calculate the fourth evaluation result. The pre-feature screening module is used to obtain third clue data based on the third evaluation results, the fourth evaluation results, and the target screening parameters; The instruction generation module is used to generate several query instructions based on third-party clue data; The large model module is used to generate corresponding query results by calling the target large model that can be queried online, based on the query command; The results generation module is used to combine the results generated by the large model module into target information filtering results.