Multi-model intelligent selection and collaborative reasoning method, device, equipment and medium

By employing a multi-model intelligent selection and collaborative reasoning method, combined with hardware performance testing and evaluation, a combination of small models with high adaptability is selected. A collaborative reasoning strategy is then used to address the limitations of browser memory and computing power, thereby improving accuracy and optimizing resources.

CN121809649APending Publication Date: 2026-04-07DRAFT (XIAMEN) INFORMATION SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Browser memory and computing power limitations prevent large AI models from running, and quantized smaller models perform poorly on complex tasks and struggle to accurately respond to user search queries.

Method used

By employing multi-model intelligent selection and collaborative reasoning methods, multimodal retrieval data is acquired, hardware performance is tested and benchmarked, performance evaluation scores are calculated, a combination of small models with high suitability is selected, and a collaborative reasoning strategy is used for task reasoning.

Benefits of technology

It effectively compensates for the insufficient accuracy of a single small model within the browser, accurately responds to user search needs, avoids resource exhaustion, and improves inference efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809649A_ABST
    Figure CN121809649A_ABST
Patent Text Reader

Abstract

The invention provides a multi-model intelligent selection and collaborative reasoning method and device, equipment and a medium. The multi-model intelligent selection and collaborative reasoning method comprises the steps that hardware performance detection and performance benchmark testing are conducted on a browser, and a performance detection score and a benchmark testing score are obtained; performing weighted summation on the performance detection score and the benchmark test score to obtain a performance evaluation score; performing model adaptation degree calculation on the plurality of target small models based on the multi-modal retrieval data and the performance evaluation scores to obtain an adaptation degree score of each target small model; performing task complexity evaluation on the multi-modal retrieval data to obtain a complexity evaluation score; according to the complexity evaluation score and the fitness score of each target small model, obtaining a small model recommendation combination; collaborative reasoning strategies corresponding to all the target small models in the small model recommendation combination are obtained according to the remaining used resources of the browser; and obtaining a task reasoning result by adopting a collaborative reasoning strategy. Therefore, the user search demand is accurately responded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of artificial intelligence technology, and more specifically, to a method, apparatus, device, and medium suitable for multi-model intelligent selection and collaborative reasoning. Background Technology

[0002] Users can enter the search terms they want to query in their browser, such as searching for a structural diagram of a device. After receiving the search instructions, the browser will perform a search based on the search terms.

[0003] In related technologies, due to the memory limitations (typically less than 2GB of available memory) and computing power limitations of the browser environment, large AI models cannot be directly loaded during query searches. Traditional large models (such as those with 7B or more parameters) simply cannot run in a browser and must be quantized into smaller models. Although quantized smaller models can run in a browser, due to the significant reduction in parameter size and precision, a single small model performs far worse than a large model on complex tasks, resulting in insufficient processing accuracy.

[0004] Therefore, using existing methods makes it difficult to accurately and effectively respond to user search results. Summary of the Invention

[0005] The embodiments described herein provide a multi-model intelligent selection and collaborative reasoning method, apparatus, device, and medium that overcomes the aforementioned problems.

[0006] Firstly, based on the content of this disclosure, a multi-model intelligent selection and collaborative reasoning method is provided, including: Acquire multimodal search data entered by the user's device in the browser; The browser is subjected to hardware performance testing and performance benchmark testing respectively, and the performance testing score and benchmark test score are obtained. The performance testing score and the benchmark test score are then weighted and summed to obtain the corresponding performance evaluation score of the browser. Based on the multimodal retrieval data and the performance evaluation score, the model fit is calculated for multiple target mini-models deployed in the browser to obtain the fit score for each target mini-model. The task complexity of the multimodal retrieval data is evaluated to obtain a complexity evaluation score; Based on the complexity evaluation score and the fitness score of each target mini-model, at least two target mini-models are selected from multiple target mini-models for optimization combination to obtain a recommended combination of mini-models; Based on the remaining resources used by the browser, the task collaboration of all the target small models in the small model recommendation combination is performed to obtain the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination; Using the aforementioned collaborative reasoning strategy, task reasoning is performed on the multimodal retrieval data through all the target small models in the small model recommendation combination, thereby obtaining the task reasoning result of the multimodal retrieval data.

[0007] Secondly, according to the present disclosure, a multi-model intelligent selection and collaborative reasoning device is provided, comprising: The acquisition module is used to acquire multimodal search data entered by the user's device in the browser; The first determining module is used to perform hardware performance testing and performance benchmark testing on the browser respectively, to obtain a performance test score and a benchmark test score; and to perform a weighted sum of the performance test score and the benchmark test score to obtain a performance evaluation score corresponding to the browser. The second determining module is used to calculate the model fit of multiple target mini-models deployed in the browser based on the multimodal retrieval data and the performance evaluation score, and obtain the fit score of each target mini-model. The evaluation module is used to evaluate the task complexity of the multimodal retrieval data and obtain a complexity evaluation score. The combination module is used to select at least two target small models from multiple target small models for optimization combination based on the complexity evaluation score and the fitness score of each target small model, so as to obtain a recommended combination of small models; The third determining module is used to perform task collaborative partitioning on all the target small models in the small model recommendation combination based on the remaining resources of the browser, so as to obtain the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination; The reasoning module is used to employ the collaborative reasoning strategy to perform task reasoning on the multimodal retrieval data using all the target small models in the small model recommendation combination, and obtain the task reasoning result of the multimodal retrieval data.

[0008] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the multi-model intelligent selection and collaborative reasoning method as described in any of the above embodiments.

[0009] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-model intelligent selection and collaborative reasoning method as described in any of the above embodiments.

[0010] The multi-model intelligent selection and collaborative reasoning method provided in this application embodiment obtains multimodal retrieval data input by a user device in a browser; performs hardware performance testing and performance benchmark testing on the browser to obtain performance test scores and benchmark test scores; and performs a weighted sum of the performance test scores and benchmark test scores to obtain the browser's corresponding performance evaluation score; calculates the model fit of multiple target small models deployed in the browser based on the multimodal retrieval data and performance evaluation scores to obtain the fit score of each target small model; evaluates the task complexity of the multimodal retrieval data to obtain a complexity evaluation score; selects at least two target small models from multiple target small models for optimized combination based on the complexity evaluation score and the fit score of each target small model to obtain a recommended combination of small models; performs task collaborative partitioning of all target small models in the recommended combination of small models based on the remaining resources used by the browser to obtain the collaborative reasoning strategy corresponding to all target small models in the recommended combination of small models; and uses the collaborative reasoning strategy to perform task reasoning on the multimodal retrieval data through all target small models in the recommended combination of small models to obtain the task reasoning result of the multimodal retrieval data. In this way, by integrating multiple small models deployed in the browser for collaborative reasoning, the problem of insufficient accuracy of a single small model can be effectively made up for, making it easier to accurately respond to user search needs; and by intelligently selecting multiple small models, the problem of resource exhaustion caused by running a large number of small models at the same time can be effectively avoided.

[0011] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein: Figure 1 This is a flowchart illustrating a multi-model intelligent selection and collaborative reasoning method provided in this disclosure.

[0013] Figure 2 This is a schematic diagram of the structure of a multi-model intelligent selection and collaborative reasoning device provided in this disclosure.

[0014] Figure 3 This is a schematic diagram of the structure of a computer device provided in this disclosure.

[0015] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0017] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.

[0018] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0019] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).

[0020] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).

[0021] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0022] Figure 1 This is a flowchart illustrating a multi-model intelligent selection and collaborative reasoning method provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the specific process of the multi-model intelligent selection and collaborative reasoning method includes: S110. Obtain multimodal retrieval data entered by the user device in the browser.

[0023] Multimodal retrieval data refers to the information a user wants to find from their browser. For example, a user might enter a query containing text and images, hoping to find related information. Multimodal retrieval data can be presented in various forms, such as text, images, audio, or video. By receiving multimodal retrieval data, the browser can parse and process different types of information. For instance, it can use multiple small models to analyze and reason about text, images, or other forms of data separately, and then integrate the results to provide the user with more accurate search feedback.

[0024] S120. Perform hardware performance testing and performance benchmark testing on the browser respectively to obtain the performance test score and benchmark test score; and perform a weighted sum of the performance test score and benchmark test score to obtain the corresponding performance evaluation score of the browser.

[0025] Hardware performance testing measures the browser's CPU (Central Processing Unit), memory, network capabilities, and GPU (Graphics Processing Unit). Performance benchmark testing assesses the browser's computing power. A weighted sum of the performance test scores and benchmark scores yields the browser's overall performance evaluation score.

[0026] In some embodiments, hardware performance testing and performance benchmarking are performed on the browser to obtain performance test scores and benchmark test scores, including: calculating the CPU performance score of the browser-equipped device based on the number of CPU cores and benchmark values; calculating the memory performance score of the browser-equipped device based on the limited memory and used memory; calculating the network performance score of the browser-equipped device based on the network performance; determining the performance test score based on the CPU performance score, memory performance score, network performance score, and GPU performance score; determining the corresponding computing power performance score based on the browser's computing power; and determining the benchmark test score based on the computing power performance score and GPU performance score.

[0027] The browser-based device's CPU performance score is calculated as follows: W_cpu = min(CPU core count / 4, 1.0) × min(benchmark value / 50, 1.0). The benchmark value reflects the device's actual performance under specific conditions. The memory performance score is W_memory = min((limited memory - used memory) / 1024MB, 1.0). The network performance score is W_network = (network type score + bandwidth score) / 2. Different network types and bandwidths correspond to different scores; a better network type indicates a better network environment, resulting in a higher network type score; better bandwidth results in a higher bandwidth score. The GPU performance score is W_gpu = min(WebGL test value / 512, 1.0). The WebGL test value reflects the device's ability to handle complex graphics tasks, derived from a graphics rendering performance test. The performance score W_performance = 0.35×W_cpu + 0.35×W_memory + 0.2×W_network + 0.1×W_gpu.

[0028] The computing performance score and GPU performance score can be assigned corresponding weights and then summed in a weighted manner to obtain the benchmark score. The computing performance score is determined based on the browser's computing power. For example, relevant data can be obtained by testing the browser with a series of computing tasks (such as mathematical operations, data processing, and code execution efficiency). Different weights are assigned to the results of different computing tasks, and the final computing performance score is obtained by weighted summation.

[0029] Therefore, by comprehensively evaluating the performance of various aspects of the device on which the browser is installed, and analyzing the performance test scores and benchmark test scores, the performance bottlenecks of the device can be identified, providing a reliable performance basis for multi-model intelligent selection and collaborative reasoning.

[0030] S130. Based on multimodal retrieval data and performance evaluation scores, calculate the model fit of multiple target mini-models deployed in the browser to obtain the fit score of each target mini-model.

[0031] These target sub-models include image recognition models, natural language processing models, and speech recognition models. By analyzing multimodal retrieval data and combining it with device performance evaluation scores, the adaptability of each sub-model in a specific environment can be further quantified. For example, adaptability calculation can comprehensively consider multiple dimensions such as model inference speed, resource utilization, and task completion accuracy, thereby generating a score reflecting the model's adaptability.

[0032] In some embodiments, model fit calculations are performed on multiple target mini-models deployed in the browser based on multimodal retrieval data and performance evaluation scores to obtain a fit score for each target mini-model. This includes: calculating the inference accuracy score corresponding to each target mini-model using the historical inference accuracy and task accuracy requirement coefficient of each target mini-model; calculating the inference speed score corresponding to each target mini-model using the expected inference time of each target mini-model and the user's expected time in the multimodal retrieval data; estimating the available memory capacity of the browser using the performance evaluation score; and calculating the inference resource score corresponding to each target mini-model based on the available memory capacity and the model memory usage of each target mini-model. The fit score of each target mini-model is determined by combining the inference accuracy score, inference speed score, inference resource score, and user satisfaction score.

[0033] The inference accuracy score W_precision is calculated as follows: W_precision = historical inference accuracy × task accuracy requirement coefficient. The task accuracy requirement coefficient can be determined by the specific application scenario and user needs, used to adjust and reflect the accuracy requirements of different tasks. The inference speed score W_speed = exp(-expected inference time ms / user-expected time ms). The inference resource score W_resource = exp(-model memory usage MB / (available memory MB × 0.8)). Available memory capacity can be estimated through performance evaluation scores, for example, by combining historical data and real-time monitoring information to accurately estimate available memory capacity. The user satisfaction score W_history = exponential moving average. The adaptability score for each target small model is W_adaptability = 0.3 × W_precision + 0.25 × W_speed + 0.3 × W_resource + 0.15 × W_history.

[0034] Therefore, based on the fitness score of each target mini-model, it is easy to select a target mini-model with a higher fitness score for inference tasks, which can improve inference efficiency while meeting the requirements of different application scenarios for accuracy, speed and resource consumption.

[0035] S140. Evaluate the task complexity of the multimodal retrieval data and obtain a complexity evaluation score.

[0036] The task complexity assessment involves analyzing multimodal retrieval data and comprehensively considering factors such as data scale, feature dimensions, and task type. The resulting complexity assessment score reflects the computational requirements and resource consumption of the task, providing a basis for subsequent model selection and inference optimization.

[0037] In some embodiments, task complexity assessment is performed on the multimodal retrieval data to obtain a complexity assessment score, including: obtaining the text length and image resolution in the multimodal retrieval data; and performing complexity assessment on the text length and image resolution in the multimodal retrieval data respectively to obtain a complexity assessment score.

[0038] The complexity assessment of text length can be achieved by calculating the ratio of the number of text characters to a preset character count threshold. The complexity assessment of image resolution can be measured based on the ratio of the total number of image pixels to a preset pixel threshold. These text length and image resolution complexity assessments effectively quantify the complexity of multimodal data during processing. Furthermore, the complexity assessment scores can be dynamically adjusted according to the needs of actual application scenarios to ensure the accuracy and applicability of the assessment results.

[0039] S150. Based on the complexity evaluation score and the fitness score of each target small model, select at least two target small models from multiple target small models for optimization combination to obtain the recommended combination of small models.

[0040] When optimizing and combining small models, the complexity evaluation score and the fitness score can be considered together. By setting weight coefficients, the influence of the complexity evaluation score and the fitness score can be balanced to ensure that the recommended combination achieves the optimal balance between performance and efficiency.

[0041] S160. Based on the remaining browser resources, perform task collaboration partitioning on all target small models in the small model recommendation combination to obtain the collaborative reasoning strategies corresponding to all target small models in the small model recommendation combination.

[0042] The collaborative inference strategies can include cascaded collaboration and parallel inference. Cascaded collaboration is driven by model confidence, refining the cascaded model relationships step by step, so that all the cascaded sub-models in the final optimized model combination can meet the confidence requirements, thereby improving the accuracy of multi-model inference. Parallel inference involves executing multiple sets of sub-models in parallel, and the execution results are weighted and summed.

[0043] In some embodiments, task collaboration is performed on all target small models in the small model recommendation combination based on the remaining resources of the browser to obtain the collaborative inference strategy corresponding to all target small models in the small model recommendation combination. This includes: if the available memory capacity of the browser is less than a preset capacity threshold, then the collaborative inference strategy corresponding to all target small models in the small model recommendation combination is determined to be cascaded collaboration; if the available memory capacity of the browser is greater than or equal to the preset capacity threshold, and the complexity evaluation score is greater than the preset complexity threshold, then the collaborative inference strategy corresponding to all target small models in the small model recommendation combination is determined to be parallel inference.

[0044] For example, if the browser's available memory capacity is less than 200MB, it can be determined that the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is cascaded collaboration. Thus, through cascaded collaboration, the execution order and dependencies of small models are optimized step by step, ensuring that the reasoning task can still be completed effectively under resource constraints. This not only reduces memory usage but also improves overall reasoning efficiency, avoiding task failure or performance degradation due to insufficient resources.

[0045] If the browser's available memory is ≥200MB and the complexity evaluation score is greater than 0.8, then the collaborative inference strategy corresponding to all target small models in the small model recommendation combination can be determined to be parallel inference. Therefore, by using parallel inference, the browser's available memory and computing resources are fully utilized, significantly improving inference speed and efficiency. The target small models in the small model recommendation combination can run simultaneously, reducing overall inference time and effectively addressing the needs of high-complexity tasks.

[0046] S170. A collaborative reasoning strategy is adopted, and all target small models in the small model recommendation combination are used to perform task reasoning on the multimodal retrieval data to obtain the task reasoning result of the multimodal retrieval data.

[0047] Among them, the collaborative reasoning strategy can dynamically adjust the execution order and dependencies of the target small model according to the characteristics of multimodal retrieval data, thereby ensuring that the reasoning process has high adaptability and stability in different scenarios.

[0048] In some embodiments, a collaborative reasoning strategy is employed to perform task reasoning on multimodal retrieval data using all target small models in the small model recommendation combination, thereby obtaining the task reasoning result for the multimodal retrieval data. This includes: if the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is cascaded collaborative, then based on the confidence of all target small models in the small model recommendation combination, multiple target small models are collaboratively combined to obtain a first model group matching the multimodal retrieval data; and collaborative reasoning is performed on the multimodal retrieval data according to the first model group to obtain the task reasoning result for the multimodal retrieval data; if the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is parallel reasoning, then functional combination is performed on all target small models in the small model recommendation combination to obtain multiple second model groups matching the multimodal retrieval data; and collaborative reasoning is performed on the multimodal retrieval data according to each second model group to obtain the task reasoning result for the multimodal retrieval data.

[0049] For example, a small model recommendation combination includes Model 1, Model 2, Model 3, and Model 4. Completing the model inference task requires the combination of multiple models. Corresponding to the cascaded collaborative strategy, Model 1 and Model 2 have the same function, Model 3 and Model 4 have the same function, and the inference task can be completed by combining Model 1 / Model 2 and Model 3 / Model 4. The resulting model groups include: Model 1 + Model 3, Model 2 + Model 3, Model 1 + Model 4, and Model 2 + Model 4. However, the confidence level of Model 1 is higher than that of Model 2, and the confidence level of Model 3 is higher than that of Model 4. Therefore, the final first model group is Model 1 + Model 2. Collaborative inference is then performed using Model 1 + Model 2 on the multimodal retrieval data to obtain the task inference result for the multimodal retrieval data. Corresponding to the parallel inference strategy, the second model group consists of: Model 1 + Model 3, Model 2 + Model 3, Model 1 + Model 4, and Model 2 + Model 4. The inference operations of Model 1 + Model 3, Model 2 + Model 3, Model 1 + Model 4, and Model 2 + Model 4 are executed in parallel to obtain the first inference result, the second inference result, the third inference result, and the fourth inference result. The first inference result, the second inference result, the third inference result, and the fourth inference result are then weighted and summed to obtain the task inference result of multimodal retrieval data.

[0050] During the weighting process, the weighting coefficients corresponding to the first, second, third, and fourth inference results can be determined based on the model weights used (e.g., the weights of multiple models are quantized, ensuring that the sum of all weighting coefficients is 1). The weights of each model are: final_weight = base_weight × (1 + confidence_bonus) × complexity_factor; base_weigh represents the model's historical accuracy, confidence_bonus = current confidence × 0.3, and complexity_factor is determined based on task complexity: complexity_factor = 1.2 when task complexity is greater than 0.7, and complexity_factor = 1.0 when task complexity is ≤ 0.7. Additionally, the model weights can be normalized: normalized_weight = final_weight / Σ(all_final_weights).

[0051] In this embodiment, multimodal retrieval data input by the user device in the browser is acquired; hardware performance testing and performance benchmark testing are performed on the browser to obtain performance test scores and benchmark test scores; the performance test scores and benchmark test scores are weighted and summed to obtain the browser's corresponding performance evaluation score; model fit calculation is performed on multiple target small models deployed in the browser based on the multimodal retrieval data and performance evaluation scores to obtain the fit score of each target small model; task complexity evaluation is performed on the multimodal retrieval data to obtain a complexity evaluation score; based on the complexity evaluation score and the fit score of each target small model, at least two target small models are selected from multiple target small models for optimization combination to obtain a recommended combination of small models; task collaborative partitioning is performed on all target small models in the recommended combination of small models based on the browser's remaining resources to obtain the collaborative reasoning strategy corresponding to all target small models in the recommended combination of small models; using the collaborative reasoning strategy, task reasoning is performed on the multimodal retrieval data through all target small models in the recommended combination to obtain the task reasoning result of the multimodal retrieval data. In this way, by integrating multiple small models deployed in the browser for collaborative reasoning, the problem of insufficient accuracy of a single small model can be effectively made up for, making it easier to accurately respond to user search needs; and by intelligently selecting multiple small models, the problem of resource exhaustion caused by running a large number of small models at the same time can be effectively avoided.

[0052] In some embodiments, the method further includes: collecting dynamic resource data from the browser during the task inference process; updating the historical inference accuracy and expected inference time of the corresponding target mini-model based on the dynamic resource data; obtaining user feedback data on the task inference results; and updating the user satisfaction score of the corresponding target mini-model based on the user feedback data.

[0053] During task inference, the system can monitor its status in real time, including intelligent model caching (LRU + frequency + accuracy priority), real-time memory monitoring (85% warning, 95% critical), adaptive resource allocation (dynamic adjustment of concurrency), and anomaly recovery mechanisms (cache cleanup → reduced concurrency → single model → GC). Based on monitoring data, adjustments are made to historical inference accuracy, expected inference time, and user satisfaction scores to ensure the validity of data used in model fit calculations.

[0054] In summary, this embodiment utilizes multiple small models for collaborative inference to compensate for the insufficient accuracy of a single small model, resulting in a significant improvement in accuracy. It adjusts concurrency strategies and model selection in real-time based on device performance, ensuring better performance across various devices and achieving dynamic load balancing. Data is processed entirely locally without uploading to a server, thoroughly protecting user privacy. There is no network latency, and the inference speed using quantized small models is faster. It can still function normally in offline environments. There are no server computing costs, and costs do not increase linearly as the user base expands.

[0055] Figure 2 This is a schematic diagram of a multi-model intelligent selection and collaborative reasoning device provided in this embodiment. The multi-model intelligent selection and collaborative reasoning device may include: The acquisition module 210 is used to acquire multimodal retrieval data entered by the user device in the browser.

[0056] The first determining module 220 is used to perform hardware performance testing and performance benchmark testing on the browser respectively, to obtain performance test scores and benchmark test scores; and to perform a weighted sum of the performance test scores and benchmark test scores to obtain the corresponding performance evaluation score of the browser.

[0057] The second determining module 230 is used to calculate the model fit of multiple target mini-models deployed in the browser based on multimodal retrieval data and performance evaluation scores, and obtain the fit score of each target mini-model.

[0058] Evaluation module 240 is used to evaluate the task complexity of multimodal retrieval data and obtain a complexity evaluation score.

[0059] The combination module 250 is used to select at least two target mini-models from multiple target mini-models for optimization combination based on the complexity evaluation score and the fitness score of each target mini-model, so as to obtain the recommended combination of mini-models.

[0060] The third determining module 260 is used to perform task collaborative partitioning on all target small models in the small model recommendation combination based on the remaining resources used by the browser, and to obtain the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination.

[0061] The reasoning module 270 is used to perform task reasoning on multimodal retrieval data by employing a collaborative reasoning strategy and using all target small models in the small model recommendation combination to obtain the task reasoning result of the multimodal retrieval data.

[0062] In this embodiment, optionally, the first determining module 220 is specifically used for: The browser's CPU performance score is calculated based on the number of CPU cores and benchmark test values. The memory performance score is calculated based on the browser's limited and used memory. The network performance score is calculated based on the browser's network performance. A performance test score is determined based on the CPU, memory, network, and GPU performance scores. A computing power score is determined based on the browser's computing capabilities. Finally, a benchmark test score is determined based on the computing power and GPU performance scores.

[0063] In this embodiment, optionally, the second determining module 230 is specifically used for: Calculate the inference accuracy score for each target mini-model by using its historical inference accuracy and task accuracy requirement coefficient; calculate the inference speed score for each target mini-model by using its expected inference time and the user's expected time in the multimodal retrieval data; estimate the browser's available memory capacity using the performance evaluation score; calculate the inference resource score for each target mini-model based on the available memory capacity and the model's memory usage; and determine the fit score for each target mini-model based on its inference accuracy score, inference speed score, inference resource score, and user satisfaction score.

[0064] In this embodiment, optionally, the evaluation module 240 is specifically used for: Obtain the text length and image resolution from the multimodal retrieval data; evaluate the complexity of the text length and image resolution from the multimodal retrieval data respectively, and obtain the complexity evaluation score.

[0065] In this embodiment, optionally, the third determining module 260 is specifically used for: If the browser's available memory capacity is less than a preset capacity threshold, then the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is determined to be cascaded collaboration; if the browser's available memory capacity is greater than or equal to the preset capacity threshold, and the complexity evaluation score is greater than the preset complexity threshold, then the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is determined to be parallel reasoning.

[0066] In this embodiment, optionally, the inference module 270 is specifically used for: If the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is cascaded collaboration, then based on the confidence of all target small models in the small model recommendation combination, multiple target small models are collaboratively combined to obtain a first model group matching the multimodal retrieval data; and collaborative reasoning is performed on the multimodal retrieval data according to the first model group to obtain the task reasoning result of the multimodal retrieval data. If the collaborative reasoning strategy corresponding to all target small models in the small model recommendation combination is parallel reasoning, then all target small models in the small model recommendation combination are functionally combined to obtain multiple second model groups matching the multimodal retrieval data; and collaborative reasoning is performed on the multimodal retrieval data according to each second model group to obtain the task reasoning result of the multimodal retrieval data.

[0067] In this embodiment, optionally, an update module is also included.

[0068] The update module is used to collect dynamic resource data from the browser during the task inference process; update the historical inference accuracy and expected inference time of the corresponding target small model based on the dynamic resource data; obtain user feedback data on the task inference results; and update the user satisfaction score of the corresponding target small model based on the user feedback data.

[0069] The multi-model intelligent selection and collaborative reasoning device provided in this disclosure can execute the above-described method embodiments. Its specific implementation principle and technical effects can be found in the above-described method embodiments, and will not be repeated here.

[0070] This application also provides a computer device. Please refer to the following for details. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0071] The computer device includes a memory 310 and a processor 320 that are interconnected via a system bus. It should be noted that only a computer device with memory 310 and processor 320 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0072] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0073] The memory 310 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, or flash card equipped on the computer device. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory 310 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method described above. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.

[0074] Processor 320 is typically used to perform overall operations of a computer device. In this embodiment, memory 310 is used to store program code or instructions, including computer operation instructions, and processor 320 is used to execute the program code or instructions stored in memory 310 or process data, such as program code that runs the methods described above.

[0075] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0076] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.

[0077] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.

[0078] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.

[0079] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0080] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0081] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0082] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" as described in this application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several units of these means may be embodied by the same item of hardware. The use of "first," "second," and "third," etc., does not indicate any order and these words should be interpreted as names. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.

[0083] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A multi-model intelligent selection and collaborative reasoning method, characterized in that, include: Acquire multimodal search data entered by the user's device in the browser; The browser is subjected to hardware performance testing and performance benchmark testing respectively, and the performance testing score and benchmark test score are obtained. The performance testing score and the benchmark test score are then weighted and summed to obtain the corresponding performance evaluation score of the browser. Based on the multimodal retrieval data and the performance evaluation score, the model fit is calculated for multiple target mini-models deployed in the browser to obtain the fit score for each target mini-model. The task complexity of the multimodal retrieval data is evaluated to obtain a complexity evaluation score; Based on the complexity evaluation score and the fitness score of each target mini-model, at least two target mini-models are selected from multiple target mini-models for optimization combination to obtain a recommended combination of mini-models; Based on the remaining resources used by the browser, the task collaboration of all the target small models in the small model recommendation combination is performed to obtain the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination; Using the aforementioned collaborative reasoning strategy, task reasoning is performed on the multimodal retrieval data through all the target small models in the small model recommendation combination, thereby obtaining the task reasoning result of the multimodal retrieval data.

2. The method according to claim 1, characterized in that, The process of performing hardware performance testing and performance benchmark testing on the browser to obtain performance test scores and benchmark test scores includes: The CPU performance score of the browser-equipped device is calculated based on the number of CPU cores and benchmark test values; the memory performance score is calculated based on the limited memory and used memory of the browser-equipped device; the network performance score is calculated based on the network performance of the browser-equipped device; and the performance test score is determined based on the CPU performance score, the memory performance score, the network performance score, and the GPU performance score. The corresponding computing power performance score is determined based on the browser's computing power; and the benchmark test score is determined based on the computing power performance score and the GPU performance score.

3. The method according to claim 1, characterized in that, The step of calculating the model fit of multiple target mini-models deployed in the browser based on the multimodal retrieval data and the performance evaluation score, to obtain a fit score for each target mini-model, includes: Calculate the inference accuracy score for each target mini-model by using the historical inference accuracy and task accuracy requirement coefficient of each target mini-model; Calculate the inference speed score corresponding to each target mini-model based on the expected inference time of each target mini-model and the user's expected time in the multimodal retrieval data; The available memory capacity of the browser is estimated based on the performance evaluation score; and the inference resource score corresponding to each target mini-model is calculated based on the available memory capacity and the memory occupied by each target mini-model. The fit score of each target mini-model is determined based on the inference accuracy score, inference speed score, inference resource score, and user satisfaction score corresponding to each target mini-model.

4. The method according to claim 1, characterized in that, The process of evaluating the task complexity of the multimodal retrieval data to obtain a complexity evaluation score includes: Obtain the text length and image resolution from the multimodal retrieval data; The complexity of the text length and image resolution in the multimodal retrieval data is evaluated separately to obtain the complexity evaluation score.

5. The method according to claim 3, characterized in that, The step of performing task collaborative partitioning on all target small models in the small model recommendation combination based on the remaining resources used by the browser to obtain the collaborative inference strategy corresponding to all target small models in the small model recommendation combination includes: If the available memory capacity of the browser is less than a preset capacity threshold, then the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination is determined to be cascaded collaboration. If the browser's available memory capacity is greater than or equal to the preset capacity threshold, and the complexity evaluation score is greater than the preset complexity threshold, then the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination is determined to be parallel reasoning.

6. The method according to claim 1, characterized in that, The method employs the collaborative reasoning strategy, using all the target small models in the small model recommendation combination to perform task reasoning on the multimodal retrieval data, to obtain the task reasoning result of the multimodal retrieval data, including: If the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination is cascaded collaboration, then based on the confidence of all the target small models in the small model recommendation combination, multiple target small models are collaboratively combined to obtain a first model group that matches the multimodal retrieval data; and collaborative reasoning is performed on the multimodal retrieval data according to the first model group to obtain the task reasoning result of the multimodal retrieval data. If the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination is parallel reasoning, then all the target small models in the small model recommendation combination are functionally combined to obtain multiple second model groups that match the multimodal retrieval data; and collaborative reasoning is performed on the multimodal retrieval data according to each second model group to obtain the task reasoning result of the multimodal retrieval data.

7. The method according to claim 3, characterized in that, Also includes: The dynamic resource data of the browser during the statistical task inference process; And update the historical inference accuracy and expected inference time of the corresponding target small model based on the dynamic resource data; Obtain user feedback data on the task inference results; and update the user satisfaction score of the corresponding target mini-model based on the user feedback data.

8. A multi-model intelligent selection and collaborative reasoning device, characterized in that, include: The acquisition module is used to acquire multimodal search data entered by the user's device in the browser; The first determining module is used to perform hardware performance testing and performance benchmark testing on the browser respectively, and obtain performance testing score and benchmark test score; The performance test score and the benchmark test score are then weighted and summed to obtain the browser's corresponding performance evaluation score. The second determining module is used to calculate the model fit of multiple target mini-models deployed in the browser based on the multimodal retrieval data and the performance evaluation score, and obtain the fit score of each target mini-model. The evaluation module is used to evaluate the task complexity of the multimodal retrieval data and obtain a complexity evaluation score. The combination module is used to select at least two target small models from multiple target small models for optimization combination based on the complexity evaluation score and the fitness score of each target small model, so as to obtain a recommended combination of small models; The third determining module is used to perform task collaborative partitioning on all the target small models in the small model recommendation combination based on the remaining resources of the browser, so as to obtain the collaborative reasoning strategy corresponding to all the target small models in the small model recommendation combination; The reasoning module is used to employ the collaborative reasoning strategy to perform task reasoning on the multimodal retrieval data using all the target small models in the small model recommendation combination, and obtain the task reasoning result of the multimodal retrieval data.

9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the multi-model intelligent selection and collaborative reasoning method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-model intelligent selection and collaborative reasoning method as described in any one of claims 1 to 7.