Vision-based large model target scene algorithm issuing method and device and computer equipment

By using a large visual model to dynamically match scene algorithms, the accuracy problem of algorithms in different environments is solved, achieving efficient adaptation and accuracy in different scenarios.

CN120614503BActive Publication Date: 2026-01-23E SURFING VISION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114925.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2026-01-23
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

In existing technologies, the application scenarios of algorithms are limited, and the accuracy is low due to environmental factors, especially when switching between indoor and outdoor scenes.

Method used

Devices based on large visual models acquire current and previous tag sets, match the tag data set with a preset correspondence list, and dynamically distribute target scene algorithms to improve scene adaptability.

Benefits of technology

It effectively avoids interference from environmental factors, improves the accuracy and recall of the algorithm under scene transformation, and reduces the false positive rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614503B_ABST
    Figure CN120614503B_ABST
Patent Text Reader

Abstract

The application relates to a target scene algorithm issuing method and device based on a visual large model and a computer device, wherein the method comprises the following steps: obtaining a current label set and a last analysis label set based on a device with a built-in visual large model; the device is associated with various scene algorithms; determining a label data set of the device according to a sub-label in the current label set and a sub-label in the last analysis label set; matching the label data set and a preset first corresponding relationship list to obtain a matching result; and issuing a target scene algorithm according to the matching result; the first corresponding relationship list has a mapping relationship between a scene algorithm and a label. Through the application, the problem of limited algorithm use scene in the related art is solved, different scene algorithms are issued for different scene algorithms, the interference of environmental factors of the scene algorithm can be effectively avoided, the accuracy of the scene algorithm issuing under scene conversion is improved, and the recall rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, apparatus and computer equipment for delivering target scene algorithms based on large visual models. Background Technology

[0002] With the development of artificial intelligence technology, algorithm capabilities have been widely applied across various industries. Taking camera algorithm applications as an example, based on video frame extraction and image frame analysis technologies, algorithms such as those for fire alerts, mask-wearing alerts, personnel intrusion monitoring, and vehicle alerts have been developed and are widely used in business monitoring and management in commerce, industry, and transportation. However, in practical applications, the use cases for these algorithms are limited. For instance, most industry algorithms are currently only applicable to one specific scenario. Algorithms trained for outdoor scenarios will experience a significant drop in performance metrics when applied to indoor scenarios, and vice versa. Furthermore, algorithms trained using both indoor and outdoor training data are less effective than those trained for a single scenario. For example, an algorithm trained for alerting electric vehicles in an elevator scenario will experience a decrease in accuracy and recall when alerting in corridors or doorways due to environmental and lighting factors.

[0003] There is currently no effective solution to the problem that the algorithm can only be used in certain scenarios on the terminal and that the accuracy of algorithm delivery is low due to environmental factors. Summary of the Invention

[0004] This embodiment provides a method, apparatus, and computer device for delivering target scene algorithms based on a large visual model, in order to solve the problems in related technologies where the application scenarios of algorithms in terminals are limited and the accuracy of algorithm delivery is low due to environmental factors.

[0005] Firstly, this embodiment provides a method for distributing target scene algorithms based on a large visual model, including:

[0006] Based on a device with a built-in large visual model, the current tag set and the previous analysis tag set are obtained; the device is associated with various scene algorithms.

[0007] The tag data set of the device is determined based on the sub-tags in the current tag set and the sub-tags in the previous analyzed tag set;

[0008] The matching results are obtained by matching the tag data set with the preset first correspondence list; and the target scene algorithm is delivered based on the matching results; the first correspondence list contains the mapping relationship between the scene algorithm and the sub-tag.

[0009] In some embodiments, the current tag set is obtained based on a device with a built-in large visual model, including:

[0010] Video streams are acquired using devices with built-in large visual models, and frames are extracted from the video streams at preset time intervals.

[0011] The image frames are subjected to image frame label analysis using a large visual model, and the current label set of the image frames is output in set form.

[0012] In some embodiments, the tag data set of the device is determined based on the current tag set and the previous analyzed tag set, including:

[0013] Compare the sub-tags in the current tag set with the sub-tags in the previous analysis tag set to determine if there are any differing sub-tags;

[0014] If there are differing sub-labels, then the sub-labels in the previous analysis label set will be reset to the sub-labels in the current label set;

[0015] If no different sub-labels exist, the device's label data set is updated based on the current label set.

[0016] In some embodiments, matching is performed based on the tag data set and a preset first correspondence list to obtain matching results, including:

[0017] Based on the tag data set and the preset first correspondence list, determine the scene algorithm corresponding to each sub-tag;

[0018] If multiple sub-labels correspond to a single scene algorithm, then the corresponding scene algorithm is determined as the target scene algorithm that the device should use.

[0019] If multiple sub-labels correspond to multiple scene algorithms, then according to the preset scoring formula, the score of each corresponding scene algorithm is calculated, and the scene algorithm with the highest score is determined as the target scene algorithm that the device should use.

[0020] If multiple sub-labels do not correspond to a scene algorithm, then the general scene algorithm is determined as the target scene algorithm that the device should use.

[0021] In some of these embodiments, the score formula is weight coefficient k1 × label 1 + weight coefficient k2 × label 2 + ... + weight coefficient kn × label n.

[0022] In some embodiments, the method further includes:

[0023] Based on the tag data set and the preset second correspondence list, the business product corresponding to the tag data set in the second correspondence list is determined, and the determined business product is marked as the target business product of the device; the second correspondence list has a mapping relationship between business products and sub-tags.

[0024] In some embodiments, the method further includes:

[0025] Promote the target business product to the users corresponding to the device.

[0026] Secondly, this embodiment provides a target scene algorithm distribution device based on a large visual model, including: an acquisition module, a processing module, and a distribution module;

[0027] The acquisition module is used to acquire the current tag set and the previous analysis tag set based on a device with a built-in large visual model; the device is associated with various scene algorithms.

[0028] The processing module is used to determine the tag data set of the device based on the sub-tags in the current tag set and the sub-tags in the previous analyzed tag set;

[0029] The delivery module is used to perform matching based on the tag data set, the scene algorithm, and a preset first correspondence list to obtain a matching result; and to complete the delivery of the target scene algorithm based on the matching result; the first correspondence list contains the mapping relationship between the scene algorithm and the sub-tag.

[0030] Thirdly, this embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the target scene algorithm delivery method based on the visual large model described in the first aspect.

[0031] Fourthly, this embodiment provides a storage medium storing a computer program that, when executed by a processor, implements the target scene algorithm delivery method based on a large visual model described in the first aspect.

[0032] Compared with related technologies, the target scene algorithm delivery method, apparatus, and computer device based on a visual large model provided in this embodiment obtain the current tag set and the previous analysis tag set through a device with a built-in visual large model; the device is associated with various scene algorithms; the tag data set of the device is determined according to the sub-tags in the current tag set and the sub-tags in the previous analysis tag set; the tag data set is matched with a preset first correspondence list to obtain the matching result; and the target scene algorithm is delivered based on the matching result; the first correspondence list has a mapping relationship between scene algorithms and tags, which solves the problem of limited algorithm usage scenarios in terminals in related technologies. By dynamically matching the corresponding scene algorithm according to the tag, different scene algorithms can be delivered for different scene algorithms, which can effectively avoid the interference of scene algorithms by environmental factors such as weather and light, improve the accuracy of scene algorithm delivery under scene transformation, and reduce the recall rate.

[0033] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0034] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0035] Figure 1 This is a hardware structure block diagram of a terminal device for a target scene algorithm distribution method based on a large visual model provided in an embodiment of this application;

[0036] Figure 2 This is a flowchart of a target scene algorithm distribution method based on a large visual model provided in an embodiment of this application;

[0037] Figure 3 This is a flowchart of step S220;

[0038] Figure 4 This is a flowchart of step S230;

[0039] Figure 5 This is a flowchart of a target scene algorithm distribution method based on a large visual model provided in another embodiment of this application;

[0040] Figure 6 This is a structural block diagram of a target scene algorithm distribution device based on a large visual model provided in an embodiment of this application.

[0041] In the diagram: 102, processor; 104, memory; 106, transmission device; 108, input / output device; 210, acquisition module; 220, processing module; 230, distribution module. Detailed Implementation

[0042] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0043] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.

[0044] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of the terminal for the target scene algorithm distribution method based on a large visual model in this embodiment. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.

[0045] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the target scene algorithm distribution method based on the visual large model in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0046] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0047] This embodiment provides a method for distributing target scene algorithms based on a large visual model. Figure 2 This is a flowchart of the target scene algorithm distribution method based on a large visual model in this embodiment, as shown below. Figure 2 As shown, the process includes the following steps:

[0048] Step S210: Based on the device with a built-in large visual model, obtain the current label set and the previous analysis label set; the device is associated with various scene algorithms;

[0049] Step S220: Determine the device's tag data set based on the sub-tags in the current tag set and the sub-tags in the previous analyzed tag set;

[0050] Step S230: Match the tag data set with the preset first correspondence list to obtain the matching result; and complete the distribution of the target scene algorithm based on the matching result; the first correspondence list contains the mapping relationship between scene algorithms and tags.

[0051] Specifically, "devices" refers to cameras, camcorders, etc., that have built-in large-scale visual models and various scene algorithms; these can also be considered the aforementioned terminals or edge devices. Large-scale visual models (Vision Foundation Models) are deep learning models trained on massive amounts of data, possessing powerful general visual understanding capabilities. They typically have a huge number of parameters (over a billion) and can handle various visual tasks, including but not limited to multimodal large-scale visual models, pure visual large-scale models, visual generative models, and industry-specific large-scale visual models. Scene algorithms include, but are not limited to, fire recognition algorithms (outdoor), fire recognition algorithms (indoor), fire recognition algorithms (forest), fire recognition algorithms (general), handwriting recognition, and face recognition. Obtaining the current label set and the previous analysis label set is a prerequisite. The previous analysis label set refers to the label data set obtained after processing the current label set from the previous analysis; it can also be the label set obtained by extracting and processing the previous video frames, without restriction. The current label set refers to the set of sub-labels obtained after processing the video frames captured by the device using the large-scale visual model. A video frame is the smallest slice of a video frame, i.e., a frozen image of the video.

[0052] The device's tag data set is obtained by comparing the sub-tags in the current tag set with the sub-tags in the previous tag set. The comparison between the sub-tags in the previous tag set and the current tag set is used to verify whether the device's screen has changed, thereby increasing the reliability of the device's tag data set and avoiding inaccurate tags caused by environmental factors such as weather, lighting, and moving objects. This reduces the occurrence of false alarms and the problem of AI illusions.

[0053] The first correspondence list is pre-set and contains a mapping relationship between scene algorithms and tags. The tag dataset is then matched against the first correspondence list to obtain the matching result. The matching result represents the target scene algorithm corresponding to the device, which is then sent to the device. This improves the accuracy of scene algorithm delivery during scene transitions and reduces the recall rate.

[0054] In related technologies, the application scenarios of algorithms in devices are limited. For example, most industry algorithms can only be applied to one type of scenario. For instance, an algorithm trained in an outdoor scenario will experience a significant drop in various algorithm metrics when applied in an indoor scenario, and vice versa. Furthermore, the performance of algorithms trained using both indoor and outdoor training data is not as good as that of algorithms trained for a single scenario because they need to accommodate different scenarios. For example, an algorithm for delivering electric vehicles, trained for use in elevator scenarios, will experience a decrease in accuracy and recall when delivered in corridors or doorways due to environmental factors such as lighting. In this embodiment, based on a device with a built-in large visual model, the current tag set and the previous analysis tag set are obtained; the device is associated with various scene algorithms; the tag data set of the device is determined according to the sub-tags in the current tag set and the sub-tags in the previous analysis tag set; the tag data set is matched with a preset first correspondence list to obtain the matching result; and the target scene algorithm is delivered based on the matching result; the first correspondence list has a mapping relationship between scene algorithms and tags, which solves the problem of limited algorithm usage scenarios in terminals in related technologies. The corresponding scene algorithm is dynamically matched according to the tag, thereby realizing the delivery of different scene algorithms for different scene algorithms. This can effectively avoid the interference of scene algorithms by environmental factors such as weather and light, improve the accuracy of scene algorithm delivery under scene transformation, and reduce the recall rate.

[0055] The steps described above are explained in detail below:

[0056] In some embodiments, the current tag set is obtained based on a device with a built-in large visual model, including the following steps:

[0057] Video streams are captured using devices with built-in large visual models, and frames are extracted from the video stream at preset time intervals.

[0058] Perform frame label analysis on the visual large model of the image frame, and output the current label set of the image frame in the form of a set.

[0059] Specifically, before extracting the video frames, the device can be initialized and a backend database can be established. The backend database is used to store various types of data in the above method embodiments, including but not limited to video streams, video frames, current tag set and previous analysis tag set, tag data set, first correspondence list and second correspondence list, etc.

[0060] After initializing the device, it acquires a video stream and then extracts frames M from the video stream at preset time intervals T. The time interval T and the number of frames M can be set according to the usage scenario and are not restricted. Frame M is input into a large-scale visual model for frame label analysis, and the current label set P of the frame is output as a set. For example, the current label set P is: image M {label 1, label 2, label 3}.

[0061] This embodiment enables the rapid and accurate extraction of image frames to complete the acquisition of the current tag set.

[0062] In some of these embodiments, such as Figure 3 As shown, step S220, which involves determining the device's tag data set based on the current tag set and the previous tag set, includes the following steps:

[0063] Step S221: Compare the sub-tags in the current tag set with the sub-tags in the previous analyzed tag set to determine whether there are any differing sub-tags;

[0064] Step S222: If there are different sub-labels, reset the sub-labels in the previous analysis label set to the sub-labels in the current label set;

[0065] Step S223: If there are no different sub-labels, update the device's label data set based on the current label set.

[0066] This embodiment can be considered as the steps of tag credibility verification and device tagging, specifically:

[0067] Compare the sub-labels in the current label set P with the sub-labels in the previous analysis label set P', traverse all sub-labels in the two sets, and determine whether there are any differing sub-labels;

[0068] If the sub-labels in the current label set P are not completely consistent with the sub-labels in the previous analyzed label set P', it is considered that there are differing sub-labels. In this case, the sub-labels in the previous analyzed label set P' are reset to the sub-labels in the current label set P, and the device's label data set is updated based on the reset current label set. The existence of differing sub-labels can be attributed to a change in the device's current usage scenario (changes in location, lighting, or screen display, etc.). For example, if the device was originally installed indoors, and the current label set P is {Label 1, Label 2, Label 3}, and algorithm A is used for face recognition; but now it has moved outdoors, and the previous analyzed label set P' is {Label 1, Label 2}, then there are differing sub-labels. If algorithm A is still used, the recognition performance will significantly decline. Therefore, subsequent steps are needed to trigger the corresponding scenario algorithm (algorithm B).

[0069] If the sub-labels in the current label set P are completely identical to the sub-labels in the previous analyzed label set P', it is assumed that there are no differing sub-labels. The device's current usage scenario (location change, lighting change, or screen change, etc.) remains unchanged. Therefore, the device's label data set L is updated based on the current label set P, making L equal to the current label set P. In this case, the previous scenario algorithm can be directly deployed to the device. For example, if the device was originally installed indoors, and the current label set P is {label 1, label 2, label 3}, then algorithm A is called for face recognition. If the device is still indoors, and the previous analyzed label set P' is {label 1, label 2, label 3}, then there are no differing sub-labels, and algorithm A is directly deployed, significantly saving computational and scheduling resources.

[0070] This embodiment utilizes the process of credibility verification and device tagging to increase the reliability of the tag data set. By tagging and storing sub-tags in the previously analyzed tag set, it avoids inaccurate tagging caused by environmental factors such as weather, light, and moving objects affecting the device, greatly reducing false alarms and solving the AI ​​illusion problem.

[0071] In some of these embodiments, such as Figure 4 As shown, step S230, which involves matching the tag data set with the preset first correspondence list to obtain the matching result, includes the following steps:

[0072] Step S231: Determine the scene algorithm corresponding to each sub-label based on the label data set and the preset first correspondence list;

[0073] Step S232: If multiple sub-labels correspond to a scene algorithm, then the corresponding scene algorithm is determined as the target scene algorithm that the device should use;

[0074] Step S233: If multiple sub-labels correspond to multiple scene algorithms, calculate the score of each corresponding scene algorithm according to the preset score formula, and determine the scene algorithm with the highest score as the target scene algorithm that the device should use.

[0075] Step S234: If multiple sub-labels do not correspond to a scene algorithm, then the general scene algorithm is determined as the target scene algorithm that the device should use.

[0076] Specifically, the first correspondence list can be as shown in Table 1 (the mapping relationship shown in Table 1 is only for illustration and does not impose any restrictions on the above mapping relationship).

[0077] Table 1

[0078]

[0079] Match the tags in the tag dataset with the first correspondence list.

[0080] If multiple sub-labels correspond to a single scenario algorithm, then the corresponding scenario algorithm is determined as the target scenario algorithm that the device should use. For example, if the labels in the label dataset are Label 1 and Label 2, and they correspond to the fire detection algorithm (outdoor), then the fire detection algorithm (outdoor) is determined as the target scenario algorithm that the device should use, and the fire detection algorithm (outdoor) is distributed for fire detection.

[0081] If multiple sub-labels correspond to multiple scene algorithms, the score for each corresponding scene algorithm is calculated according to a preset scoring formula, and the scene algorithm with the highest score is determined as the target scene algorithm to be used by the device. The scoring formula is: weight coefficient k1 × label 1 + weight coefficient k2 × label 2 + ... + weight coefficient kn × label n. For example, if the labels in the label dataset are label 1, label 2, and label 5; assuming each label has a score of 1; weight coefficient k1 is 0.1; weight coefficient k2 is 0.1; weight coefficient k5 is 0.3; then the scene algorithm corresponding to label 1 and label 2 has a score of 0.2; and the scene algorithm corresponding to label 5 has a score of 0.3. Therefore, corresponding to the fire detection algorithm (forest), the fire detection algorithm (forest) is determined as the target scene algorithm to be used by the device, and the fire detection algorithm (forest) is distributed for fire detection. The weight coefficients and the score of label 2 can be set according to the usage scenario and are not limited thereto. In other embodiments, pre-trained neural network models or other methods can also be used to calculate the score of the scene algorithm.

[0082] If multiple sub-tags do not correspond to a specific scenario algorithm, then the general scenario algorithm will be determined as the target scenario algorithm that the device should use. For example, if the tag in the tag dataset is tag 7, then it corresponds to the fire detection algorithm (general). Therefore, the fire detection algorithm (general) will be determined as the target scenario algorithm that the device should use, and the fire detection algorithm (general) will be distributed for fire detection.

[0083] This embodiment delivers corresponding scene algorithms based on the device's tag data set, ensuring that the scene algorithms adapt to the usage environment. For multiple tags, different weight coefficients for each tag can be used to score and sort them, selecting the optimal scene algorithm for scheduling and delivery. By combining scene algorithms with the usage environment, interference from environmental factors is avoided, significantly improving the accuracy of scene algorithm delivery during scene transitions and reducing recall.

[0084] In some of these embodiments, such as Figure 5 As shown, the target scene algorithm delivery method based on the large visual model also includes the following steps:

[0085] Step S240: Based on the tag data set and the preset second correspondence list, determine the business product corresponding to the tag data set in the second correspondence list, and mark the determined business product as the target business product of the device; the second correspondence list has a mapping relationship between business products and sub-tags.

[0086] Specifically, step S240 can be considered to be executed after S230; the second correspondence list has the mapping relationship between business products and tags, as shown in Table 2 (the mapping relationship shown in Table 2 is only for illustration and does not impose any restrictions on the above mapping relationship).

[0087] Table 2

[0088]

[0089] The tags in the tag dataset are matched with the second correspondence list, and then the business product corresponding to the tag dataset in the second correspondence list is determined based on the matching results; the determined business product is then marked as the target business product for the device. Each device has user attributes, allowing identification of which user it belongs to.

[0090] Example: Taking user A's device as an example. It has been determined that the tags in user A's device tag data set are {tag1, tag2, tag3}, and user A's confirmed business product is {fire detection}. Matching user A's tags with the second correspondence list, we obtain user A's current business product as {fire detection, intrusion detection, specific action detection (smoking, etc.)}. Therefore, user A's target business product is updated to {fire detection, intrusion detection, specific action detection}. At this point, user A can be considered a potential customer for fire detection, intrusion detection, and specific action detection. Potential customers refer to a group of potential clients.

[0091] This embodiment enables the processing of potential user attribute tags, combining the device usage environment with the user's potential application products, and converging the user's potential needs into clear and specific product capabilities.

[0092] In some embodiments, the target scene algorithm delivery method based on a large visual model further includes the following steps:

[0093] Promote the target business product to the users corresponding to the device.

[0094] Specifically, since the users associated with the devices have been identified as potential customers for the corresponding products, promotions can be made to these users based on these tags, thereby increasing conversion rates and reducing marketing costs.

[0095] It should be noted that the steps shown in the above process or in the flowchart of the accompanying figures can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0096] This embodiment also provides a target scene algorithm distribution device based on a large visual model. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below can refer to combinations of software and / or hardware that perform a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0097] Figure 6 This is a structural block diagram of the target scene algorithm distribution device based on a large visual model in this embodiment, as shown below. Figure 6 As shown, the device includes: an acquisition module 210, a processing module 220, and a distribution module 230;

[0098] The acquisition module 210 is used to acquire the current label set and the previous analysis label set based on a device with a built-in large visual model; the device is associated with various scene algorithms.

[0099] Processing module 220 is used to determine the tag data set of the device based on the sub-tags in the current tag set and the sub-tags in the previous analyzed tag set;

[0100] The distribution module 230 is used to match the tag data set with the preset first correspondence list to obtain the matching result; and to distribute the target scene algorithm according to the matching result; the first correspondence list contains the mapping relationship between scene algorithm and tag.

[0101] The aforementioned device solves the problem of limited application scenarios for algorithms in terminals in related technologies. It dynamically matches the corresponding scenario algorithm based on the tag, thereby enabling the delivery of different scenario algorithms for different scenarios. This effectively avoids interference from environmental factors such as weather and lighting, improves the accuracy of scenario algorithm delivery during scenario transitions, and reduces the recall rate.

[0102] In some embodiments, the acquisition module 210 is also used to acquire video streams based on a device with a built-in visual large model and extract frame images from the video stream at preset time intervals.

[0103] Perform frame label analysis on the visual large model of the image frame, and output the current label set of the image frame in the form of a set.

[0104] In some embodiments, the processing module 220 is further configured to compare the sub-tags in the current tag set with the sub-tags in the previous analyzed tag set to determine whether there are any differing sub-tags;

[0105] If there are differing sub-labels, the sub-labels in the previous analysis label set will be reset to the sub-labels in the current label set;

[0106] If no different sub-labels exist, the device's label data set will be updated based on the current label set.

[0107] In some embodiments, the distribution module 230 is further configured to determine the scene algorithm corresponding to each sub-tag based on the tag data set and the preset first correspondence list;

[0108] If multiple sub-labels correspond to a single scene algorithm, then the corresponding scene algorithm is determined as the target scene algorithm that the device should use.

[0109] If multiple sub-labels correspond to multiple scene algorithms, then according to the preset scoring formula, the score of each corresponding scene algorithm is calculated, and the scene algorithm with the highest score is determined as the target scene algorithm that the device should use.

[0110] If multiple sub-labels do not correspond to a scene algorithm, then the general scene algorithm will be determined as the target scene algorithm that the device should use.

[0111] The recommended module score formula is: weight coefficient k1 × tag 1 + weight coefficient k2 × tag 2 + ... + weight coefficient kn × tag n.

[0112] In some embodiments, the target scene algorithm distribution device based on the large visual model further includes a promotion module;

[0113] The promotion module is used to determine the business products corresponding to the tag data set in the second correspondence list based on the tag data set and the preset second correspondence list, and to mark the determined business products as the target business products of the device; the second correspondence list contains the mapping relationship between business products and sub-tags.

[0114] In some embodiments, the promotion module is also used to promote the target business product to the users corresponding to the device.

[0115] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0116] This embodiment also provides a computer device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0117] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0118] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0119] S1, based on a device with a built-in large visual model, obtains the current tag set and the previous analysis tag set; the device is associated with various scene algorithms;

[0120] S2, Based on the sub-labels in the current label set and the sub-labels in the previous analyzed label set, determine the device's label data set;

[0121] S3, perform matching based on the tag data set and the preset first correspondence list to obtain the matching result; and complete the distribution of the target scene algorithm based on the matching result; the first correspondence list contains the mapping relationship between scene algorithms and tags.

[0122] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.

[0123] Furthermore, in conjunction with the target scene algorithm delivery method based on a large visual model provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the target scene algorithm delivery methods based on a large visual model in the above embodiments.

[0124] It should be noted that all information and data involved in this application are authorized by the user or fully authorized by all parties and will be used legally.

[0125] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0126] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.

[0127] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0128] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A method for distributing target scene algorithms based on a large visual model, characterized in that, include: Based on a device with a built-in large visual model, the current tag set and the previous analysis tag set are obtained; the device is associated with various scene algorithms. Based on the sub-tags in the current tag set and the sub-tags in the previous analyzed tag set, the tag data set of the device is determined, which includes: Compare the sub-tags in the current tag set with the sub-tags in the previous analysis tag set to determine if there are any differing sub-tags; If a difference sub-label exists, the sub-labels in the previous analysis label set are reset to the sub-labels in the current label set. The label data set is matched with a preset first correspondence list to obtain a matching result. The target scene algorithm is then deployed based on the matching result. The first correspondence list contains the mapping relationship between the scene algorithm and the sub-labels. The existence of a difference sub-label refers to changes in the location, lighting, or screen of the current usage scene of the device. If no difference sub-label exists, the device's label data set is updated based on the current label set, and the previous target scene algorithm is directly issued; where no difference sub-label exists means that the location, lighting, or screen of the current usage scene of the device has not changed; the device's label data set is verified for credibility based on whether the screen of the device has changed.

2. The target scene algorithm distribution method based on a large visual model according to claim 1, characterized in that, Based on devices with built-in large visual models, the current tag set is obtained, including: Video streams are acquired using devices with built-in large visual models, and frames are extracted from the video streams at preset time intervals. The image frames are subjected to image frame label analysis using a large visual model, and the current label set of the image frames is output in set form.

3. The target scene algorithm distribution method based on a large visual model according to claim 1 or claim 2, characterized in that, The matching results are obtained by matching the tag data set with the preset first correspondence list, including: Based on the tag data set and the preset first correspondence list, determine the scene algorithm corresponding to each sub-tag; If multiple sub-labels correspond to a single scene algorithm, then the corresponding scene algorithm is determined as the target scene algorithm that the device should use. If multiple sub-labels correspond to multiple scene algorithms, then according to the preset scoring formula, the score of each corresponding scene algorithm is calculated, and the scene algorithm with the highest score is determined as the target scene algorithm that the device should use. If multiple sub-labels do not correspond to a scene algorithm, then the general scene algorithm is determined as the target scene algorithm that the device should use.

4. The target scene algorithm distribution method based on a large visual model according to claim 3, characterized in that, The scoring formula is: weight coefficient k1 × label 1 + weight coefficient k2 × label 2 + ... + weight coefficient kn × label n.

5. The target scene algorithm distribution method based on a large visual model according to claim 3, characterized in that, The method further includes: Based on the tag data set and the preset second correspondence list, the business product corresponding to the tag data set in the second correspondence list is determined, and the determined business product is marked as the target business product of the device; the second correspondence list has a mapping relationship between business products and sub-tags.

6. The target scene algorithm distribution method based on a large visual model according to claim 5, characterized in that, The method further includes: Promote the target business product to the users corresponding to the device.

7. A target scene algorithm distribution device based on a large visual model, characterized in that, include: The module includes an acquisition module, a processing module, and a distribution module. The acquisition module is used to acquire the current tag set and the previous analysis tag set based on a device with a built-in large visual model; the device is associated with various scene algorithms. The processing module is used to determine the tag data set of the device based on the sub-tags in the current tag set and the sub-tags in the previous analyzed tag set, which includes: Compare the sub-tags in the current tag set with the sub-tags in the previous analysis tag set to determine if there are any differing sub-tags; If a difference sub-tag exists, the sub-tags in the previously analyzed tag set are reset to the sub-tags in the current tag set. Based on the distribution module, matching is performed according to the tag data set, the scene algorithm, and a preset first correspondence list to obtain a matching result. The target scene algorithm is then distributed based on the matching result. The first correspondence list contains the mapping relationship between the scene algorithm and the sub-tags. The existence of a difference sub-tag refers to changes in the location, lighting, or screen of the current usage scene of the device. If no difference sub-label exists, the device's label data set is updated based on the current label set, and the previous target scene algorithm is directly issued based on the issuing module; wherein, no difference sub-label means that the location, lighting or screen of the current usage scene of the device has not changed; the device's label data set is verified for credibility based on whether the screen of the device has changed.

8. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the target scene algorithm delivery method based on any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the target scene algorithm delivery method based on the visual large model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for recommending product

    CN108242016A

  • Continual selection of scenarios based on identified tags describing contextual environment of a user for execution by an artificial intelligence model of the user by an autonomous personal companion

    CN111201539A