Video identification method and device, electronic equipment and storage medium
By building an intelligent distribution system with multiple recognition terminals, shopping videos are dynamically allocated and secondary recognition is triggered when the first recognition fails. This solves the problem of low video recognition accuracy in intelligent unmanned retail vending machines, improves system stability and user experience, and reduces operating costs.
Patent Information
- Application Number
- CN202511436822.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-02-10
AI Technical Summary
The existing intelligent unmanned retail vending machines have low video recognition accuracy, and are particularly difficult to adapt to complex and special shopping scenarios, resulting in low recognition accuracy and poor system stability, which increases the company's operating costs and the risk of product damage.
By building an intelligent distribution system with multiple recognition terminals, the rules engine first dynamically allocates shopping videos to the first recognition terminal for recognition. If the first recognition result is abnormal, the rules engine is triggered to schedule the second recognition terminal for a second recognition. The final result is determined by arbitrating the comparison of the two recognition results, thus optimizing the settlement process.
It improved the accuracy of video recognition, enhanced system stability, shortened recognition time, reduced labor costs, optimized user experience, reduced cargo damage, and increased total transaction volume.
Smart Images

Figure CN121505501A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video recognition, and in particular to a video recognition method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the rapid development of the new retail model, unmanned retail has become an important trend in the industry, and intelligent unmanned retail cabinets have been widely used and promoted in many cities. Through the combination of video recognition and automatic settlement, this technology has significantly improved the convenience and diversity of user shopping.
[0003] However, existing intelligent unmanned retail cabinets generally rely on a single video recognition service provider to identify and settle user shopping behavior. For complex shopping scenarios (such as stacked goods, occlusion, similar goods differentiation, fast picking actions, etc.) or special shopping behaviors (such as returning goods, replacing goods, etc.), the identification algorithm of a single identification provider is difficult to adapt, resulting in low accuracy of video recognition. SUMMARY
[0004] The main purpose of the embodiments of the present application is to provide a video recognition method, device, electronic device, and storage medium, which can solve the problem of low accuracy of video recognition of intelligent unmanned retail cabinets in the prior art and improve the accuracy of video recognition when users shop through intelligent unmanned retail.
[0005] To achieve the above purpose, a first aspect of the embodiments of the present application provides a video recognition method applied to a distribution system, the method comprising: capturing a user shopping process through a camera arranged on a cabinet to obtain a shopping video; determining a first identification end from a plurality of identification ends; performing first identification of user shopping behavior in the shopping video through the first identification end to obtain a first identification result; if the first identification result meets a preset condition, determining a second identification end from the plurality of identification ends, wherein the second identification end is different from the first identification end, and the preset condition includes that the first identification result indicates that user shopping is not detected to be successful, or, in the case where the first identification result includes an identification duration, the identification duration is not within a preset duration range, the identification duration being a duration for which the first identification end identifies user shopping behavior in the shopping video; performing second identification of user shopping behavior in the shopping video through the second identification end to obtain a second identification result; analyzing the first identification result and the second identification result to determine a target identification result.
[0006] In some embodiments, the first recognition end is determined from the plurality of recognition ends, comprising: a plurality of first candidate recognition ends are selected from the plurality of recognition ends, the first candidate recognition end being a recognition end whose recognition accuracy in a first historical period is greater than a first preset threshold, and / or the candidate recognition end being a recognition end whose recognition time length spent in the first historical period is less than a second preset threshold; a first target candidate recognition end is selected from the plurality of first candidate recognition ends, the first target candidate recognition end having the highest recognition accuracy in the first historical period; if the number of shopping videos to be processed by the first target candidate recognition end is less than a first preset number, the first target candidate recognition end is taken as the first recognition end; if the number of shopping videos to be processed by the first target candidate recognition end is greater than or equal to the first preset number, a first other recognition end is taken as the plurality of first candidate recognition ends, and the step of selecting the first target candidate recognition end from the plurality of first candidate recognition ends is executed, the first other recognition end being a recognition end other than the first target candidate recognition end among the plurality of first candidate recognition ends.
[0007] In some embodiments, the first recognition end is determined from the plurality of recognition ends, comprising: if the number of shopping videos to be processed by each of the first candidate recognition ends is greater than or equal to the first preset number, a second target candidate recognition end is selected from the plurality of first candidate recognition ends, the second target candidate recognition end having the shortest recognition time length spent in the first historical period; if the number of shopping videos to be processed by the second target candidate recognition end is less than a second preset number, the second target candidate recognition end is taken as the first recognition end; if the number of shopping videos to be processed by the second target candidate recognition end is greater than or equal to the second preset number, a second other recognition end is taken as the plurality of first candidate recognition ends, and the step of selecting the second target candidate recognition end from the plurality of first candidate recognition ends is executed, the second other recognition end being a recognition end other than the second target candidate recognition end among the plurality of first candidate recognition ends.
[0008] In some embodiments, the second recognition end is determined from the plurality of recognition ends, comprising: a plurality of second candidate recognition ends are determined from the plurality of recognition ends, the second candidate recognition end having a number of times of being selected for secondary recognition in a first historical period greater than a third preset threshold; a target recognition result obtained by each of the second candidate recognition ends in the first historical period is obtained; selecting the second recognition terminal from a plurality of second candidate recognition terminals, wherein the second recognition terminal satisfies at least one of the following: the target recognition result of the second recognition terminal indicates that the number of times of detecting user shopping success is the largest; the target recognition result of the second recognition terminal indicates that the ratio of the number of times of detecting user shopping success to the number of times of being selected for secondary recognition in the first historical period is the largest.
[0009] In some embodiments, in the case that the first recognition result indicates that user shopping success is not detected, the first recognition result includes that the user does not take the goods in the cabinet, or the user puts goods that do not belong to the cabinet; the determining of the second recognition terminal from a plurality of recognition terminals includes: obtaining target recognition results obtained by a plurality of third candidate recognition terminals for secondary recognition in a first historical period, wherein the plurality of third candidate recognition terminals are recognition terminals other than the first recognition terminal in the plurality of recognition terminals; selecting the second recognition terminal from a plurality of third candidate recognition terminals, wherein the second recognition terminal satisfies at least one of the following: in the case that the first recognition result obtained by the first-time recognition in the first historical period includes that the user does not take the goods in the cabinet, the target recognition result obtained by being selected for secondary recognition indicates that the number of times of detecting user shopping success is the largest; or, in the case that the first recognition result obtained by the first-time recognition in the first historical period includes that the user puts goods that do not belong to the cabinet, the target recognition result obtained by being selected for secondary recognition indicates that the number of times of detecting user shopping success is the largest.
[0010] In some embodiments, the analyzing of the first recognition result and the second recognition result to determine a target recognition result includes: obtaining the number of goods categories and the number of goods in the first recognition result, and the number of goods categories and the number of goods in the second recognition result; comparing the number of goods categories in the first recognition result and the number of goods categories in the second recognition result, and taking the recognition result with more number of goods categories as the target recognition result; if the number of goods categories in the first recognition result and the number of goods categories in the second recognition result are the same, taking the recognition result with the largest number of goods in the first recognition result and the second recognition result as the target recognition result.
[0011] In some embodiments, after the analyzing the first recognition result and the second recognition result, and determining the target recognition result, the method further includes: obtaining a commodity type and a corresponding commodity quantity in the target recognition result; performing commodity settlement according to the commodity type and the corresponding commodity quantity.
[0012] To achieve the above object, a second aspect of the embodiments of the present application provides a video recognition device, applied to a distribution system, the device comprising: a shooting module, configured to shoot a user shopping process through a camera arranged on a cabinet to obtain a shopping video; a first determination module, configured to determine a first recognition terminal from a plurality of recognition terminals; a first recognition module, configured to perform first recognition on a user shopping behavior in the shopping video through the first recognition terminal to obtain a first recognition result; a second determination module, configured to determine a second recognition terminal from the plurality of recognition terminals if the first recognition result meets a preset condition, wherein the second recognition terminal is different from the first recognition terminal, and the preset condition includes that the first recognition result indicates that the user shopping is not successful, or, in a case where the first recognition result includes a recognition time length, the recognition time length is not located in a preset time length range, the recognition time length being a time length for the first recognition terminal to recognize the user shopping behavior in the shopping video; a second recognition module, configured to perform second recognition on the user shopping behavior in the shopping video through the second recognition terminal to obtain a second recognition result; an analysis module, configured to analyze the first recognition result and the second recognition result to determine a target recognition result.
[0013] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method of the first aspect when executing the computer program.
[0014] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer readable storage medium, storing a computer program, and the computer program implementing the method of the first aspect when executed by a processor.
[0015] The video recognition method, device, electronic equipment and storage medium provided by the application, through constructing an intelligent distribution system containing a rule engine and scene linkage, first distribute the shopping video to the first recognition end based on the historical performance data for recognition; if the first recognition result is abnormal (such as recognition timeout or returning non-shopping result), the rule engine is automatically triggered to schedule the second recognition end for secondary recognition; the system compares the two recognition results, arbitrates the final recognition result according to the preset rule (such as preferentially selecting the result of identifying more commodity categories), and completes commodity settlement accordingly, thereby effectively improving the recognition accuracy and system reliability, reducing the loss and labor cost. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a flowchart of the video recognition method provided by the embodiment of the application; Figure 2 is a flowchart of the multi-scene automatic distribution recognition system provided by the embodiment of the application; Figure 3 is an interaction diagram of the automatic distribution system provided by the embodiment of the application; Figure 4 is a video recognition diagram provided by the embodiment of the application; Figure 5 is a rule engine configuration diagram provided by the embodiment of the application; Figure 6 is a scene linkage configuration diagram provided by the embodiment of the application; Figure 7 is a structural diagram of the video recognition device provided by the embodiment of the application; Figure 8 is a hardware structure diagram of the electronic equipment provided by the embodiment of the application. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical scheme and advantages of the application more clear, the application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application, and are not used to limit the application.
[0018] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.
[0020] With the continuous evolution of new retail formats, unmanned retail mode has gradually become the development trend of the industry. As the core carrier of unmanned retail, intelligent unmanned retail cabinets have been widely applied and promoted in many cities in China, greatly improving the convenience of user shopping and enriching the diversity of user consumption choices. However, with the rapid expansion of the intelligent unmanned retail industry, the outstanding problems existing in the shopping video processing technology of the existing intelligent unmanned retail cabinet have gradually been exposed, mainly focusing on two core dimensions: first, the recognition accuracy of user shopping video is low, and second, the overall stability of the shopping video recognition system is insufficient. The above problems have become key bottlenecks restricting the further development of the industry.
[0021] In the prior art, after the user completes shopping, the shopping video generated by the cabinet is usually only distributed to a single identification merchant for analysis and identification, and the system directly completes the settlement operation of the user's order according to the identification result output by the single identification merchant. This mode has significant defects in actual application: for complex shopping scenarios (such as the user taking multiple different categories of goods at the same time, the goods being placed and shielding each other) or special shopping scenarios (such as the user temporarily putting the goods back into the cabinet, the local part of the shopping video being blurred or insufficient light), the identification algorithm of the single identification merchant is difficult to adapt, resulting in that the recognition accuracy cannot be effectively guaranteed, and thus a large number of abnormal orders are generated; such abnormal orders need to rely on manual customer service intervention for checking and processing, greatly increasing the operating cost of the enterprise.
[0022] At the same time, the existing system has a strong dependence on a single identification merchant. When the identification merchant has technical failures (such as server downtime, abnormal identification algorithm), network interruption or temporary service interruption, etc., the system lacks effective alternative identification channels and can only passively wait for the identification merchant to resume service, resulting in that the identification time of the shopping video is completely uncontrollable, and the order settlement process is stalled. In addition, due to the hard rules of mainstream payment channels such as WeChat and Alipay (i.e. the number of in-transit orders under the same user account cannot exceed a preset threshold), if the previous shopping order is in a "processing" state due to identification stagnation, the user cannot initiate a new shopping request, resulting in that the user cannot open the cabinet door again for shopping, which seriously damages the continuity of user consumption.
[0023] The above problems of low recognition accuracy and poor system stability will directly or indirectly cause a series of chain negative effects such as the increase of enterprise goods loss risk, the decrease of gross merchandise volume (GMV), the increase of customer complaint rate, and the increase of manual customer service cost, etc.
[0024] Based on this, the embodiment of the present application provides a video recognition method and device, electronic equipment and storage medium, aiming to provide a technical solution which can effectively improve the user shopping video recognition accuracy, enhance the system stability, shorten the recognition time, reduce the labor cost and optimize the user experience. The embodiment of the present application sets different automatic distribution rules and linkage scenes, monitors the user shopping video recognition time and different recognition results in real time, and once the preset threshold is reached, the distribution rule engine is triggered, the corresponding linkage scene is automatically switched to, other identification companies are called for identification, and then the recognition results of different manufacturers are compared, and settlement is made according to the rule setting.
[0025] The video recognition method, device, electronic equipment and storage medium provided by the embodiment of the present application are specifically described through the following embodiments. First, the video recognition method in the embodiment of the present application is described.
[0026] The video recognition method provided by the embodiment of the present application relates to the technical field of video recognition. The video recognition method provided by the embodiment of the present application can be applied to a terminal, can be applied to a server end, and can also be software running in a terminal or a server end. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc. The server end can be configured as an independent physical server, can be configured as a server cluster or a distributed system composed of multiple physical servers, can also be configured as a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and basic cloud computing services such as big data and artificial intelligence platforms, etc. The software can be an application that implements the video recognition method, but is not limited to the above forms.
[0027] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0028] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0029] Figure 1 This is an optional flowchart of the video recognition method provided in this application embodiment, which is applied to a distribution system. Figure 1 The method may include, but is not limited to, steps S100 to S600.
[0030] Step S100: The user's shopping process is filmed by a camera installed on the shelf to obtain a shopping video.
[0031] In this embodiment, the distribution system (multi-scenario automatic distribution and recognition system) is a cloud-based software system responsible for intelligently managing and scheduling multiple video recognition services (i.e., "recognition terminals") to analyze user shopping videos and ultimately determine the basis for settlement. The distribution system can receive shopping videos uploaded from vending machines and distribute them to the recognition terminals for behavior recognition; it can define distribution rules, recognition duration thresholds, result anomaly judgment conditions, and linkage trigger logic; and it can dynamically call backup recognition terminals for secondary recognition based on the trigger conditions of the rule engine.
[0032] Specifically, the system first uses pre-installed cameras on the smart unmanned retail vending machine to record the user's entire shopping process in real time (from opening the door to entering the vending machine, taking the goods, to closing the door and leaving), creating a shopping video that records the user's shopping behavior. After recording, the vending machine preprocesses the shopping video (such as noise reduction and format adaptation to meet the parsing requirements of the recognition end) and uploads the preprocessed shopping video to the cloud distribution system in preparation for subsequent recognition.
[0033] Step S200: Determine the first identification end from multiple identification ends.
[0034] In this embodiment, the distribution system selects and determines the first identification terminal for initial identification based on a preset identification terminal distribution ratio rule from multiple candidate identification terminals (i.e., manufacturers with shopping video analysis capabilities, which can output information such as product retrieval results and abnormal behavior judgment).
[0035] Specifically, the distribution ratio rules for each recognition terminal are formulated by the distribution system's rule engine, based on historical recognition performance data for each terminal. This data originates from monthly sampling tests of each terminal (approximately 1000 shopping video recognition records are sampled from each terminal), including recognition accuracy (such as the correct recognition rate for product type / quantity and the correct recognition rate for abnormal scenes), average recognition time (the time taken from receiving the video to outputting the result in a single recognition), and the number of orders with a recognition time exceeding 60 seconds. For example, after analyzing the historical performance data, the rule engine sets the "recognition terminal with the highest recognition accuracy" to a 40% distribution ratio (i.e., 40% of shopping videos are preferentially distributed to this terminal), the "recognition terminal with the shortest average recognition time" to a 30% distribution ratio, and the "recognition terminal with superior overall performance (both accuracy and time)" to a 30% distribution ratio. The distribution system then selects the recognition terminal that meets the above distribution ratios as the first recognition terminal based on the current allocation queue of videos to be recognized.
[0036] For example, the identifier with the highest recognition accuracy (identifier) can be allocated a 40% distribution ratio, the identifier with the shortest average recognition time can be allocated a 30% distribution ratio, and the identifier with both high recognition accuracy and average recognition time can be allocated the remaining 30% distribution ratio. This embodiment can dynamically adjust the video identifier distribution ratio threshold, recognition time threshold, and recognition result category, setting different distribution ratios for different identifiers based on their specific circumstances. Simultaneously, for different recognition times, the rule engine's internal timed task continuously polls the recognition results of each shopping video. This embodiment can set the identifier distribution ratio based on monthly sampling results of each identifier's recognition accuracy and average recognition time. Each identifier has approximately 1000 samples. The sampling items include the recognition accuracy of each manufacturer's shopping videos, average recognition time, recognition accuracy of different abnormal recognition results, total recognition time, number of recognition times exceeding 60 seconds, and number of different abnormal recognition results.
[0037] Step S300: The user's shopping behavior in the shopping video is initially identified by the first identification terminal to obtain a first identification result.
[0038] In this embodiment, the distribution system sends the shopping video to a designated first identification terminal. The first identification terminal analyzes the user's shopping behavior in the video, including the type and quantity of goods taken, whether there are any abnormal behaviors (such as maliciously placing non-display items), and whether the video image is clear (such as whether it cannot be resolved due to lighting issues). After completing the initial identification, the first identification terminal generates a first identification result and sends it back to the distribution system. The first identification result is used to indicate whether the user's purchase was successful.
[0039] Specifically, the first identification result can be categorized into two types: normal results and abnormal results. Normal results include clearly defined product types and quantities (e.g., "1 bottle of cola, 2 loaves of bread"). Abnormal results include situations where a successful purchase was not detected, such as not taking the product, abnormal user behavior, abnormal placement of the item, or video anomalies (unable to be parsed). Furthermore, the first identification result may also include the duration of the shopping video identified by the first identification device.
[0040] Step S400: If the first recognition result meets the preset conditions, then a second recognition end is determined from multiple recognition ends, wherein the second recognition end is different from the first recognition end. The preset conditions include the first recognition result indicating that no successful user shopping was detected, or, if the first recognition result includes a recognition duration, the recognition duration is not within a preset duration range. The recognition duration refers to the time taken by the first recognition end to recognize the user's shopping behavior in the shopping video.
[0041] In this embodiment, the distribution system sends the first identification result to the rule engine, which then determines whether the first identification result meets the preset conditions. If it does, the system determines a second identification end that is different from the first identification end from among multiple candidate identification ends.
[0042] Specifically, the preset conditions include two types of situations: The first identification result indicates that the user's purchase was not detected. That is, the first identification result is an abnormal result such as no product taken, abnormal user behavior, abnormal placement, or abnormal video. This means that the first identification did not obtain valid shopping information and a second identification verification is required.
[0043] The first recognition result includes the recognition time, and this recognition time is not within the preset time range. The recognition time refers to the total time taken by the first recognition terminal to recognize the shopping video; the preset time range is dynamically configured by the rule engine (e.g., the default setting is 0 seconds < recognition time ≤ 20 seconds). If the recognition time is 0 seconds (the first recognition terminal does not respond) or greater than 20 seconds (the recognition delay is too high), it is determined that the preset conditions are met.
[0044] The rules for determining the second identification terminal are formulated by the scenario linkage module of the distribution system, and are selected based on the principle of adapting to the type of anomaly. If the condition is triggered because the identification time is not within the preset range, the identification terminal with the shortest historical average identification time can be selected as the second identification terminal (based on the average time data of monthly sampling); if the condition is triggered because the first identification result is an abnormal result, the identification terminal with the highest accuracy in identifying this type of abnormal result can be selected as the second identification terminal (e.g., if the first identification result is that the product was not taken, then the identification terminal with the highest accuracy in identifying the scenario of not taking the product is selected).
[0045] Step S500: The second recognition terminal performs secondary recognition on the user's shopping behavior in the shopping video to obtain a second recognition result.
[0046] In this embodiment, the distribution system's rule engine triggers scene linkage. The scheduling engine determines a second recognition endpoint from the remaining recognition endpoints based on the rules. For example, the rules might specify: when recognition endpoint A times out, automatically select the second fastest-responding recognition endpoint B as the second recognition endpoint; when recognition endpoint A returns a blurry video recognition result, automatically select the recognition endpoint C, which has the strongest ability to handle such abnormal scenarios, as the second recognition endpoint. The distribution system sends the same shopping video to the second recognition endpoint (e.g., recognition endpoint B) for secondary recognition. After completing the analysis, recognition endpoint B returns a second recognition result. This second recognition result indicates whether a successful user purchase has been detected.
[0047] Step S600: Analyze the first recognition result and the second recognition result to determine the target recognition result.
[0048] In this embodiment, the distribution system's rule engine compares and analyzes the first and second identification results, and determines the target identification result for order settlement according to preset priority rules. The priority rules are as follows: The identification result that identifies more product types is preferred. If one identification result records more product types than the other (e.g., the first result identifies 2 products, and the second result identifies 3 products), the second identification result is preferred. If the number of product types is the same, the identification result that identifies more products is preferred (e.g., both results identify 2 products, but the first result identifies 3 and the second result identifies 4). If both identification results are abnormal results indicating no successful user purchase, the second identification result is selected by default, as it is adapted to the current abnormality type and has better accuracy.
[0049] The flowchart of the multi-scenario automatic distribution and recognition system is as follows: Figure 2 As shown, this embodiment mainly includes the following steps: 1. The user opens the door to shop; 2. The vending machine uploads the shopping video to the automatic distribution system; 3. The rule engine defines the distribution ratio of the recognition merchants, the recognition time, and the recognition result comparison rules; 4. Scene linkage sends the shopping video to other recognition merchants for video recognition; 5. Settlement is made according to the recognition results.
[0050] The automatic distribution system interaction is as follows: Figure 3As shown, the system consists of video recognition, a rules engine, and scene linkage. Video recognition involves distributing user shopping videos to recognition providers for identification, and then processing the payment based on the final recognition result. The rules engine configures different thresholds for different device capabilities, analyzes the information reported by the devices, and calculates whether the information will reach the threshold. Scene linkage involves calling other recognition providers configured for scene linkage to perform recognition for those that reach the corresponding threshold, comparing the two recognition results, and selecting the final recognition result for payment.
[0051] This embodiment avoids order delays caused by relying on a single recognition terminal by switching between multiple recognition terminals (calling the second recognition terminal when the first recognition terminal malfunctions), shortens the total recognition time, ensures timely order settlement, and allows users to shop again normally, thus improving system availability and user experience. The rule engine can dynamically adjust parameters such as the distribution ratio of recognition terminals, preset duration range, and abnormal result types. The scene linkage module can update the selection rules of the second recognition terminal according to changes in the performance of the recognition terminal, adapting to the operational needs of unmanned vending machines in different regions and at different times. By introducing a secondary recognition and result comparison mechanism, especially when there is doubt about the initial recognition result, a review is initiated, which effectively corrects the misjudgment of a single recognition terminal, significantly reducing the probability of missed recognition and misrecognition, thereby reducing cargo loss and increasing GMV.
[0052] In some embodiments, step S200 may include, but is not limited to, steps S210 to S240: Step S210: Select a plurality of first candidate recognition ends from the plurality of recognition ends. The first candidate recognition ends are the recognition ends among the plurality of recognition ends whose recognition accuracy is greater than a first preset threshold in a first historical period, and / or, the candidate recognition ends are the recognition ends among the plurality of recognition ends whose recognition time spent in the first historical period is less than a second preset threshold. Step S220: Select a first target candidate identification end from multiple first candidate identification ends, wherein the first target candidate identification end has the highest identification accuracy in the first historical time period; Step S230: If the number of shopping videos to be processed by the first target candidate recognition terminal is less than the first preset number, then the first target candidate recognition terminal is used as the first recognition terminal. Step S240: If the number of shopping videos to be processed by the first target candidate recognition terminal is greater than or equal to the first preset number, then the first other recognition terminal is used as multiple first candidate recognition terminals, and the process jumps to the step of selecting the first target candidate recognition terminal from multiple first candidate recognition terminals. The first other recognition terminal is the recognition terminal other than the first target candidate recognition terminal among the multiple first candidate recognition terminals.
[0053] In this embodiment, the selection criteria for the first candidate recognition end combines historical accuracy and recognition time to ensure that the candidate recognition end has basic performance guarantees. The selection of the first target candidate recognition end prioritizes the recognition end with the highest accuracy to improve the reliability of the initial recognition. The number of videos to be processed is determined by dynamically adjusting the recognition end load through a preset threshold to avoid processing delays caused by task backlog on a single recognition end. When the load of the first target candidate recognition end is too high, the system automatically switches to other candidate recognition ends, forming a dynamic polling mechanism. The first historical time period can be defined as the most recent calendar month, that is, the performance of the recognition end is evaluated based on the sampling data of the past month, ensuring the timeliness and representativeness of the data. The first preset threshold (accuracy threshold) can be set to 90%, that is, only recognition ends with a sampling accuracy > 90% in the past month are included as candidates, excluding low-accuracy recognition ends. The second preset threshold (duration threshold) can be set to 20 seconds, that is, only recognition ends with an average recognition time < 20 seconds in the past month are included as candidates, excluding high-latency recognition ends. The first preset quantity (load threshold) can be set based on the actual concurrent processing capability of the recognition end. For example, a single recognition end can process a maximum of 50 shopping videos at the same time (to avoid delays caused by task backlog). Therefore, the first preset quantity is set to 50, that is, the number of videos to be processed is less than 50, which is considered to be controllable load.
[0054] Specifically, the system first filters candidate recognition endpoints based on historical data, selecting those with satisfactory accuracy or processing speed, forming a first set of candidate recognition endpoints. Then, it selects the recognition endpoint with the highest historical accuracy from this set as the first target candidate recognition endpoint. If the number of videos to be processed by this endpoint does not exceed a preset threshold, the task is directly assigned; if it exceeds the threshold, the endpoint is excluded, and the system reselects the next highest-accurate endpoint from the remaining candidate endpoints. For example, when the first preset number is set to 50, if the first target candidate recognition endpoint already has 50 videos to be processed, the system will skip this endpoint and select the next best candidate. This process is repeated until a recognition endpoint that meets the load conditions is found. Through dynamic load balancing, the system effectively avoids processing delays caused by endpoint overload while ensuring recognition accuracy.
[0055] This embodiment uses historical accuracy and recognition time as filtering criteria to ensure that the first candidate recognition ends are all high-performance recognition ends, reducing the probability of initial recognition errors or delays from the source and reducing the triggering of subsequent secondary recognitions. By checking the number of tasks to be processed by the recognition ends, new tasks are avoided from being assigned to overloaded recognition ends, thereby reducing recognition latency and improving system response speed. If the optimal recognition end is temporarily unavailable or overloaded, the system automatically switches to the suboptimal recognition end, ensuring the continuity of the recognition process, reducing dependence on a single recognition end, realizing dynamic adjustment and load balancing of recognition ends, and effectively improving the processing capacity and response speed of the entire video recognition system.
[0056] In some embodiments, step S200 may also include, but is not limited to, steps S250 to S270: Step S250: If the number of shopping videos to be processed by each of the first candidate recognition terminals is greater than or equal to the first preset number, then a second target candidate recognition terminal is selected from the plurality of first candidate recognition terminals, wherein the second target candidate recognition terminal has the shortest recognition time in the first historical period. Step S260: If the number of shopping videos to be processed by the second target candidate recognition terminal is less than the second preset number, then the second target candidate recognition terminal is used as the first recognition terminal. Step S270: If the number of shopping videos to be processed by the second target candidate recognition terminal is greater than or equal to the second preset number, then the second other recognition terminal is used as multiple first candidate recognition terminals, and the process jumps to the step of selecting the second target candidate recognition terminal from multiple first candidate recognition terminals. The second other recognition terminal is the recognition terminal other than the second target candidate recognition terminal among the multiple first candidate recognition terminals.
[0057] In this embodiment, the selection criterion for the second target candidate recognition end is the shortest recognition time, which is derived from the statistical average video processing time of each recognition end in the historical period; the setting of the second preset number is dynamically adjusted based on the processing capacity of the recognition end to ensure that the selected recognition end has the ability to process new tasks in real time; the jump execution step traverses all candidate recognition ends through a loop mechanism until a recognition end that meets the load conditions is found.
[0058] Specifically, when the number of videos to be processed by all first-selection recognition terminals exceeds a first preset number, the system switches to a filtering mode prioritizing recognition duration. The system retrieves the candidate recognition terminal with the shortest recognition duration from historical data as the second target candidate recognition terminal and checks whether its current task queue is below a second preset threshold. If the condition is met, the current task is assigned to that recognition terminal; otherwise, the recognition terminal is excluded and the filtering process is re-executed until a usable recognition terminal is found. This mechanism, through a dynamic priority switching strategy, automatically activates recognition terminals prioritizing processing speed when the load on candidate recognition terminals with high accuracy is too high, ensuring timely execution of shopping video recognition tasks and avoiding system delays caused by task backlog.
[0059] This embodiment selects the recognition terminal with the shortest recognition time as the first choice, which can minimize user waiting time and improve user experience. It can also avoid assigning too many tasks to a single recognition terminal while ensuring recognition efficiency, thereby achieving a balanced distribution of recognition tasks. This dynamic allocation mechanism can also adapt to the performance fluctuations of different recognition terminals and always maintain the optimal operating state of the system.
[0060] In some embodiments, step S400 may include, but is not limited to, steps S410 to S430: Step S410: Determine multiple second candidate identification terminals from multiple identification terminals, wherein the number of times the second candidate identification terminal is selected for secondary identification in the first historical time period is greater than a third preset threshold. Step S420: Obtain the target recognition result obtained by each of the second candidate recognition ends in the first historical time period through secondary recognition; Step S430: Select the second identification end from a plurality of second candidate identification ends, wherein the second identification end satisfies at least one of the following: The target recognition result of the second recognition end indicates that the user has been detected to have made the most successful purchases. The target recognition result of the second recognition terminal indicates that the ratio of the number of times a user successfully made a purchase to the number of times the second recognition terminal was selected for secondary recognition during the first historical period is the largest.
[0061] In this embodiment, the selection of the second candidate identification terminal is based on its call frequency in historical time periods. A call frequency exceeding a preset threshold indicates that the identification terminal has a higher task allocation priority. The target identification result is obtained by retrieving the record data of each second candidate identification terminal performing secondary identification within the historical time period. The record data includes the number of times a user's shopping was successfully detected and the corresponding number of calls. The selection logic of the second identification terminal is divided into two paths: the first path prioritizes the identification terminal with the most successful times, and the second path prioritizes the identification terminal with the largest ratio of successful times to calls. The two paths can be applied independently or in combination. If multiple second candidate identification terminals simultaneously meet the above two conditions (e.g., a terminal has the highest number of successful times and the highest efficiency), it is directly determined as the second identification terminal; if different terminals meet two conditions respectively (e.g., A has the highest efficiency and B has the most successful times), the distribution system flexibly determines the second identification terminal based on the anomaly type of the current first identification. The second candidate identification terminal refers to the identification terminal that has been selected for secondary identification more than the third preset threshold in the first historical time period. The core is to select identification terminals with sufficient experience in secondary identification. The third preset threshold can be set to 300 times, meaning that only recognition terminals with more than 300 secondary recognitions in the past month will be included as candidates, ensuring that they have experience in handling various abnormal scenarios. The secondary recognition success efficiency refers to the ratio of the number of times the target recognition result indicates a successful purchase in the historical secondary recognition of the second candidate recognition terminal to the total number of times that terminal was selected for secondary recognition during the same period, which is used to quantify the effectiveness of the secondary recognition of the recognition terminal.
[0062] Specifically, when the initial identification triggers a secondary identification requirement, the system first filters out candidate identification terminals that have been called for secondary identification more than a third preset threshold within a historical period. These candidate identification terminals are considered to have processing capacity or resource stability due to their high frequency of participation in secondary identification. Subsequently, the system retrieves the result data of each candidate identification terminal performing secondary identification within the historical period, counts the number of times each candidate identification terminal successfully detected a user's successful purchase in secondary identification, and calculates the ratio of this number to its total number of calls. During the selection process, if the highest number of successful transactions is used as the criterion, the system directly selects the candidate identification terminal with the highest historical number of successful transactions as the secondary identification terminal; if the largest ratio is used as the criterion, the candidate identification terminal with the highest success detection rate is selected first. This mechanism ensures that the selection of secondary identification terminals is not only based on historical call frequency but also combined with their actual identification effect, thereby improving the accuracy of secondary identification and the utilization rate of system resources.
[0063] This embodiment improves the accuracy of secondary identification by selecting the identification end with the best secondary identification effect from historical data, reduces cargo damage caused by incorrect identification, and reduces the workload of manual verification. At the same time, by selecting the identification end with a high identification success rate, the overall identification time is shortened and the user experience is improved.
[0064] In some embodiments, if the first identification result indicates that no successful user purchase was detected, the first identification result includes whether the user did not take the goods from the shelf or whether the user placed goods that do not belong to the shelf. Step S400 may also include, but is not limited to, steps S440 to S450: Step S440: Obtain the target recognition results obtained by each of the multiple third candidate recognition ends in the first historical time period through secondary recognition. The multiple third candidate recognition ends are recognition ends other than the first recognition end among the multiple recognition ends. Step S450: Select the second identification end from the plurality of third candidate identification ends, wherein the second identification end satisfies at least one of the following: In the first historical period, if the first identification result obtained during the first identification includes the case where the user has not taken the goods from the cabinet, the target identification result obtained by selecting for secondary identification indicates that the user has been detected to have made the most successful purchases. Alternatively, if, within the first historical period, the first identification result obtained during the initial identification includes items placed by the user that do not belong to the vending machine, the target identification result obtained by selecting the second identification indicates that the user has been detected to have made the most successful purchases.
[0065] In this embodiment, the selection scope of the third candidate identification end is limited to the remaining identification ends after excluding the first identification end, ensuring that the secondary identification end and the first identification end are complementary in terms of algorithmic logic. The selection criterion for the third candidate identification end is based on its error correction capability for similar abnormal scenarios in historical data. Specifically, it is quantitatively evaluated by statistically analyzing the frequency or proportion of each identification end outputting correct results after handling similar abnormal scenarios in historical periods. For example, when the initial identification result is that the user did not take the goods, the system retrieves the identification records of each third candidate identification end handling similar abnormal scenarios in the past thirty days, calculates the number of times each identification end successfully identified the user's actual shopping behavior in such scenarios, and selects the identification end with the highest number of successful identifications as the secondary identification end.
[0066] Specifically, after the system outputs the identification result of the user not taking the goods or placing abnormal goods at the initial identification end, it automatically triggers a secondary identification process. First, the initial identification end is excluded from the identification end list, generating a set of third candidate identification ends. Then, the historical database is accessed to extract the identification records of each third candidate identification end in handling scenarios that perfectly match the current anomaly type within a preset period. For example, for the scenario of the user not taking the goods, the system filters out the number of times each identification end has output correct results when handling similar anomalies in the past. If identification end B has 85 successful times in this scenario and identification end C has 72 successful times, then identification end B is selected as the secondary identification end. Through this mechanism, the secondary identification end can provide optimal identification capabilities for specific anomaly types, effectively improving the accuracy of the target identification results and reducing the risk of misjudgment caused by the limitations of a single identification end algorithm.
[0067] This embodiment selects the most suitable second recognition end for secondary recognition based on different anomaly recognition results. This targeted selection can improve the accuracy of secondary recognition, reduce misjudgments, and thus improve the overall recognition accuracy. At the same time, by selecting the recognition end with the best historical performance, the stability and reliability of the system can be improved. Furthermore, tasks can be allocated according to the expertise of different recognition ends, making full use of the advantages of each recognition end and improving the overall efficiency of the system.
[0068] In some embodiments, step S600 may include, but is not limited to, steps S610 to S630: Step S610: Obtain the number of product types and the number of products in the first identification result, and the number of product types and the number of products in the second identification result; Step S620: Compare the number of product types in the first identification result with the number of product types in the second identification result, and take the identification result with more product types as the target identification result; Step S630: If the number of product types in the first identification result is the same as the number of product types in the second identification result, then the identification result with the largest number of products in the first identification result and the second identification result is taken as the target identification result.
[0069] In this embodiment, the number of product types refers to the total number of different product categories in the identification results, while the number of products refers to the number of items of the same product category. During the comparison process, the number of product types is compared first; if the two are the same, the number of products is then compared. For example, if the first identification result contains 3 types of products with a total quantity of 5 items, and the second identification result contains 2 types of products with a total quantity of 6 items, then the first identification result is selected as the target identification result. If both identification results contain 3 types of products, but the first identification result has a total quantity of 4 items and the second identification result has a total quantity of 5 items, then the second identification result is selected.
[0070] Specifically, after the initial and secondary identification processes, the system extracts the product type and quantity data from both identification results. By comparing the number of product types, the system prioritizes results covering more product categories to reduce errors caused by a single identification device missing specific categories. When the number of product types is the same, the system further compares the total number of products and selects the result with the larger quantity to more comprehensively reflect the amount of products actually taken by the user. For example, if the initial identification device fails to detect a small, obscured item due to algorithm limitations, but the secondary identification device identifies the item through multi-angle analysis, the system automatically selects the secondary identification result containing more product types by comparing the differences in the number of product types, thereby improving the accuracy of order settlement.
[0071] This embodiment compares the number of product types and the quantity of products in the two identification results, and selects the more comprehensive and detailed identification result as the final result. This reduces identification errors and omissions, more accurately reflects the user's actual shopping behavior, reduces the risk of product damage, and improves the accuracy of settlement. At the same time, because a comparison mechanism of two identification results is adopted, even if there is a deviation in a single identification, it can be supplemented and corrected by the result of the other identification, thereby improving the reliability and stability of the overall identification system.
[0072] In some embodiments, the steps prior to step S100 may include, but are not limited to, the following steps: Obtain the shopping request and authentication information initiated by the user; The authentication information is verified to obtain the verification result. If the verification result is successful, the locker is unlocked and the user's shopping process is recorded as a video.
[0073] In this embodiment, a shopping request refers to a user-initiated instruction to open the vending machine for shopping. This is triggered by scanning a QR code with WeChat / Alipay or by facial recognition. Users can scan the vending machine's QR code or initiate a facial recognition request in front of the vending machine's facial recognition module to complete the shopping request. The identity verification information corresponds to the triggering method of the shopping request, divided into QR code scanning and facial recognition scenarios. In the QR code scanning scenario, this involves the user's unique WeChat / Alipay account identifier (such as WeChat OpenID or Alipay UID) and account status (such as whether a valid payment method is bound or whether there are any outstanding orders). In the facial recognition scenario, this involves the user's facial image feature data (which must comply with privacy protection requirements and is only used for comparison and verification; the original image is not stored). The verification result refers to the verification conclusion of the distribution system (or the vending machine's local control system) on the identity verification information, categorized only as "passed" or "failed." Vending machine unlocking control refers to the unlocking command sent by the distribution system (or local control system) to the vending machine's lock module. Shopping video recording is triggered synchronously with the unlocking command; the recording start time must be exactly the same as the unlocking time to ensure coverage of the entire process of the user opening the door and retrieving goods.
[0074] Specifically, the vending machine receives user operations through pre-set interactive modules (QR code stickers, facial recognition cameras) and simultaneously obtains shopping requests and identity verification information. The distribution system performs multi-dimensional verification of the obtained identity verification information to ensure that the user has legitimate shopping permissions. In addition to verifying account validity, it also needs to verify whether the user currently has any incomplete orders in transit (such as previously purchased but unsettled orders). If the number of in-transit orders is greater than or equal to the payment channel limit (e.g., 1), the verification fails. If all the above verification items meet the requirements (valid account, facial recognition matching, no excessive in-transit orders), the verification result is passed; if any verification item is not met (e.g., frozen account, facial recognition mismatch, existing in-transit orders), the verification result is failed, and the reason is displayed on the vending machine screen / user APP pop-up (e.g., "You have incomplete orders, please process them first"), terminating the subsequent process. Only when the verification result is passed is the simultaneous operation of unlocking the vending machine and recording video triggered.
[0075] This embodiment employs a strict identity verification mechanism to prevent unauthorized use and reduce the risk of product damage; video recording is initiated immediately upon successful verification to avoid missing key shopping behavior data; and malicious behavior is effectively prevented by verifying the user's credit status and order history.
[0076] In some embodiments, the steps following step S600 may include, but are not limited to, the following steps: Obtain the product type and corresponding product quantity from the target identification results; The settlement of goods shall be carried out according to the types and quantities of the goods.
[0077] In this embodiment, after comparing the two identification results, the system extracts the product information recorded in the target identification result. The target identification result contains product type and quantity data confirmed after the two identification comparisons. After receiving this data, the settlement system retrieves the corresponding product price database, calculates the total amount, and generates an order. Product information includes product name and corresponding quantity. The system queries the database based on the product name to obtain the standard unit price, multiplies the unit price by the quantity to obtain the individual item amount, and adds up all individual item amounts to generate the total order amount. The total order amount is transmitted to the payment platform through an encrypted interface. The payment platform deducts the payment based on the user's pre-deposited account or bound payment method. If the deduction is successful, the system updates the inventory data and generates a transaction record; if the deduction fails, the system triggers a retry mechanism and sends a failure notification to the maintenance terminal. After settlement, the order status is updated to "completed," releasing the in-transit order limit under the user's account. The product price database is updated in real time to ensure that the settlement amount is consistent with the current pricing. After the order is generated, the container inventory synchronization mechanism is triggered to update the inventory quantity in real time.
[0078] This embodiment seamlessly integrates video recognition with transaction settlement to form a complete automated shopping process; through standardized product information parsing and price query mechanisms, it ensures the accuracy of settlement amounts; the automated settlement process reduces human intervention and significantly improves transaction processing speed.
[0079] The video recognition diagram of this embodiment is as follows: Figure 4 As shown in the diagram, the rule engine configuration is as follows: Figure 5 As shown in the diagram, the scene linkage configuration is as follows: Figure 6 As shown.
[0080] The specific steps of the multi-scenario automatic distribution system include: Step 1: The user scans the code or uses facial recognition. After successful verification, the container door is opened. Step 2: The vending machine begins recording video of the user taking the goods; Step 3: Upload the shopping video to the cloud-based multi-scenario distribution system; Step 4: For the received user shopping video, the corresponding recognition provider is called for initial recognition according to the recognition provider distribution ratio rules; Step 5: Receive the recognition result from the recognition provider's callback and send the recognition result to the rules engine; Step 6: The rule engine determines whether scene linkage is needed through two dimensions. Triggering one of them will trigger scene linkage. First, the rule engine will periodically check whether the recognition time exceeds a certain threshold. For example, if the recognition time is equal to 0 seconds or greater than 20 seconds, it means that the recognition time is abnormal and other manufacturers need to be called for recognition again through scene linkage. Second, the rule engine will determine the abnormal results based on the initial recognition result, such as: no product taken, abnormal user behavior, abnormal placement, abnormal video, etc. As long as the recognition provider returns one of the abnormal results, other recognition providers will be called for recognition again through scene linkage. Step 7: When the rule is triggered, scene linkage is performed. According to the rule configuration, video recognition requests are sent to other corresponding recognition systems, and the recognition results of the automatic distribution system are called back. Step 8: The rules engine automatically compares the recognition results of different recognition companies and determines one of the recognition results as the final settlement result; Step 9: Based on the recognition results, the user's shopping video will be used to complete the order settlement.
[0081] This application embodiment utilizes a multi-identifier dynamic collaboration mechanism to automatically trigger a secondary identification process when the initial identification fails. It combines the algorithmic advantages of different identification devices for result comparison, significantly improving identification accuracy and system fault tolerance in complex shopping scenarios. Simultaneously, it monitors identification duration and result validity in real-time under preset conditions, preventing order processing delays caused by single identification device failures. This application embodiment, by defining identification durations with different rules, enables rapid video identification and settlement, while also preventing issues arising from a single identification vendor, improving user experience and system availability. By defining different identification result rules, it allows for multiple identifications of different vendors in certain scenarios, improving accuracy, reducing product damage, and increasing GMV. Through different scenario-linked rules, it calls different identification vendors, automating video identification control, automatically comparing identification results, and processing settlements, reducing customer complaints and the cost of handling abnormal orders manually.
[0082] Please see Figure 7 This application also provides a video recognition device 700, which can implement the above-described video recognition method and is applied to a distribution system. The device includes: The shooting module 10 is used to capture the user's shopping process through a camera installed on the shelf to obtain shopping videos; The first determining module 20 is used to determine the first identifying end from multiple identifying ends; The first identification module 30 is used to perform initial identification of user shopping behavior in the shopping video through the first identification terminal to obtain a first identification result; The second determining module 40 is used to determine a second identifying end from multiple identifying ends if the first identification result meets preset conditions, wherein the second identifying end is different from the first identifying end. The preset conditions include the first identification result indicating that no successful user shopping was detected, or, if the first identification result includes an identification time, the identification time is not within a preset time range. The identification time refers to the time taken by the first identifying end to identify the user's shopping behavior in the shopping video. The second recognition module 50 is used to perform secondary recognition on the user's shopping behavior in the shopping video through the second recognition terminal to obtain a second recognition result; The analysis module 60 is used to analyze the first identification result and the second identification result to determine the target identification result.
[0083] In some implementations, the first determining module 20 may include: The first selection submodule is used to select a plurality of first candidate recognition ends from a plurality of recognition ends. The first candidate recognition ends are recognition ends whose recognition accuracy is greater than a first preset threshold in a first historical period, and / or, the candidate recognition ends are recognition ends whose recognition time spent in the first historical period is less than a second preset threshold. The second selection submodule is used to select a first target candidate recognition end from multiple first candidate recognition ends, wherein the first target candidate recognition end has the highest recognition accuracy in the first historical time period. The first judgment submodule is used to select the first target candidate recognition terminal as the first recognition terminal if the number of shopping videos to be processed by the first target candidate recognition terminal is less than the first preset number. The first jump rotor module is used to, if the number of shopping videos to be processed by the first target candidate recognition terminal is greater than or equal to the first preset number, then to select the first other recognition terminal as multiple first candidate recognition terminals and jump to the step of selecting the first target candidate recognition terminal from multiple first candidate recognition terminals. The first other recognition terminal is the recognition terminal other than the first target candidate recognition terminal among the multiple first candidate recognition terminals.
[0084] In some implementations, the first determining module 20 may further include: The third selection submodule is used to select a second target candidate recognition terminal from multiple first candidate recognition terminals if the number of shopping videos to be processed by each first candidate recognition terminal is greater than or equal to the first preset number, wherein the second target candidate recognition terminal has the shortest recognition time in the first historical period. The second judgment submodule is used to select the second target candidate recognition terminal as the first recognition terminal if the number of shopping videos to be processed by the second target candidate recognition terminal is less than the second preset number. The second jump rotor module is used to select the second target candidate recognition terminal as a plurality of first candidate recognition terminals if the number of shopping videos to be processed by the second target candidate recognition terminal is greater than or equal to the second preset number. The second other recognition terminal is a recognition terminal other than the second target candidate recognition terminal among the plurality of first candidate recognition terminals.
[0085] In some implementations, the second determining module 40 may include: The determination submodule is used to determine multiple second candidate recognition ends from multiple recognition ends, wherein the number of times the second candidate recognition end is selected for secondary recognition in the first historical time period is greater than a third preset threshold. The first acquisition submodule is used to acquire the target recognition result obtained by each of the second candidate recognition ends in the first historical time period through secondary recognition; The fourth selection submodule is used to select the second identification end from a plurality of second candidate identification ends, wherein the second identification end satisfies at least one of the following: the target identification result of the second identification end indicates that the number of times the user's shopping success is detected is the largest; the target identification result of the second identification end indicates that the ratio of the number of times the user's shopping success is detected to the number of times the second identification end is selected for secondary identification in the first historical period is the largest.
[0086] In some implementations, the second determining module 40 may further include: The second acquisition submodule is used to acquire the target recognition results obtained by each of the multiple third candidate recognition ends in the first historical time period through secondary recognition. The multiple third candidate recognition ends are recognition ends other than the first recognition end among the multiple recognition ends. The fifth selection submodule is used to select the second identification terminal from a plurality of the third candidate identification terminals, wherein the second identification terminal satisfies at least one of the following: when the first identification result obtained during the first identification in the first historical period includes the user not taking the goods from the cabinet, the target identification result obtained by selecting the second identification terminal indicates that the user has successfully made a purchase the most times; or, when the first identification result obtained during the first identification in the first historical period includes the user putting in goods that do not belong to the cabinet, the target identification result obtained by selecting the second identification terminal indicates that the user has successfully made a purchase the most times.
[0087] In some implementations, the analysis module 60 may include: The third acquisition submodule is used to acquire the number of product types and the number of products in the first identification result, as well as the number of product types and the number of products in the second identification result; The comparison submodule is used to compare the number of product types in the first identification result and the number of product types in the second identification result, and to take the identification result with more product types as the target identification result; The third judgment submodule is used to take the recognition result with the largest number of product types in the first recognition result and the second recognition result as the target recognition result if the number of product types in the first recognition result and the number of product types in the second recognition result are the same.
[0088] In some embodiments, the apparatus may further include: The acquisition module is used to acquire the product types and corresponding product quantities in the target recognition results; The settlement module is used to settle accounts for goods based on the type and quantity of the goods.
[0089] The specific implementation of this video recognition device is basically the same as the specific implementation of the video recognition method described above, and will not be repeated here.
[0090] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described video recognition method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0091] Please see Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 801 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the video recognition method of the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0092] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video recognition method.
[0093] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0094] The video recognition method, video recognition device, electronic device, and storage medium provided in this application embodiment construct an intelligent distribution system that includes a rule engine and scene linkage. First, based on historical performance data, shopping videos are dynamically allocated to a first recognition end for recognition. If the first recognition result is abnormal (such as recognition timeout or returning a result of no purchase), the rule engine is automatically triggered to schedule a second recognition end for secondary recognition. The system compares the two recognition results and arbitrates according to preset rules (such as prioritizing the result of recognizing more types of goods) to determine the final recognition result, and completes the goods settlement accordingly. This effectively improves the recognition accuracy and system reliability, and reduces goods damage and labor costs.
[0095] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0096] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0099] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0100] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0101] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0102] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0103] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0104] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0105] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A video recognition method, characterized in that, Applied to a distribution system, the method includes: The shopping process is filmed by cameras installed on the shelves to obtain shopping videos; The first identification end is determined from multiple identification ends; The first recognition device performs initial recognition of the user's shopping behavior in the shopping video to obtain a first recognition result. If the first identification result meets the preset conditions, then a second identification end is determined from multiple identification ends, wherein the second identification end is different from the first identification end. The preset conditions include the first identification result indicating that no successful user purchase was detected, or, if the first identification result includes identification time, the identification time is not within the preset time range. The identification time refers to the time taken by the first identification end to identify the user's shopping behavior in the shopping video. The second recognition device performs secondary recognition on the user's shopping behavior in the shopping video to obtain a second recognition result. The first and second identification results are analyzed to determine the target identification result.
2. The method according to claim 1, characterized in that, Determining the first identification end from multiple identification ends includes: Multiple first candidate recognition ends are selected from multiple recognition ends. The first candidate recognition ends are recognition ends whose recognition accuracy is greater than a first preset threshold in a first historical period, and / or, the candidate recognition ends are recognition ends whose recognition time spent in the first historical period is less than a second preset threshold. A first target candidate identification end is selected from multiple first candidate identification ends, and the first target candidate identification end has the highest identification accuracy in the first historical time period; If the number of shopping videos to be processed by the first target candidate recognition terminal is less than the first preset number, then the first target candidate recognition terminal will be used as the first recognition terminal. If the number of shopping videos to be processed by the first target candidate recognition terminal is greater than or equal to the first preset number, then the first other recognition terminal is used as multiple first candidate recognition terminals, and the process jumps to the step of selecting the first target candidate recognition terminal from multiple first candidate recognition terminals. The first other recognition terminal is the recognition terminal other than the first target candidate recognition terminal among the multiple first candidate recognition terminals.
3. The method according to claim 2, characterized in that, Determining the first identification end from multiple identification ends includes: If the number of shopping videos to be processed by each of the first candidate recognition terminals is greater than or equal to the first preset number, then a second target candidate recognition terminal is selected from the multiple first candidate recognition terminals, and the second target candidate recognition terminal has the shortest recognition time in the first historical period. If the number of shopping videos to be processed by the second target candidate recognition terminal is less than the second preset number, then the second target candidate recognition terminal will be used as the first recognition terminal. If the number of shopping videos to be processed by the second target candidate recognition terminal is greater than or equal to the second preset number, then the second other recognition terminal is used as multiple first candidate recognition terminals, and the process jumps to the step of selecting the second target candidate recognition terminal from multiple first candidate recognition terminals. The second other recognition terminal is the recognition terminal other than the second target candidate recognition terminal among the multiple first candidate recognition terminals.
4. The method according to claim 1, characterized in that, Determining the second identification terminal from multiple identification terminals includes: Multiple second candidate identification ends are determined from multiple identification ends, and the number of times the second candidate identification end is selected for secondary identification in the first historical period is greater than a third preset threshold. Obtain the target recognition result obtained by performing secondary recognition on each of the second candidate recognition ends within the first historical time period; The second identification end is selected from a plurality of second candidate identification ends, wherein the second identification end satisfies at least one of the following: The target recognition result of the second recognition end indicates that the user has been detected to have made the most successful purchases. The target recognition result of the second recognition terminal indicates that the ratio of the number of times a user successfully made a purchase to the number of times the second recognition terminal was selected for secondary recognition during the first historical period is the largest.
5. The method according to claim 1, characterized in that, If the first identification result indicates that the user's purchase was not successfully detected, the first identification result includes the user not taking the goods from the cabinet or the user putting in goods that do not belong to the cabinet; Determining the second identification terminal from multiple identification terminals includes: The target recognition results obtained by each of the multiple third candidate recognition ends performing secondary recognition within the first historical time period are obtained, wherein the multiple third candidate recognition ends are recognition ends other than the first recognition end among the multiple recognition ends; The second identification end is selected from a plurality of the third candidate identification ends, wherein the second identification end satisfies at least one of the following: In the first historical period, if the first identification result obtained during the first identification includes the case where the user has not taken the goods from the cabinet, the target identification result obtained by selecting for secondary identification indicates that the user has been detected to have made the most successful purchases. Alternatively, if, within the first historical period, the first identification result obtained during the initial identification includes items placed by the user that do not belong to the vending machine, the target identification result obtained by selecting the second identification indicates that the user has been detected to have made the most successful purchases.
6. The method according to claim 1, characterized in that, The step of analyzing the first identification result and the second identification result to determine the target identification result includes: Obtain the number of product types and the number of products in the first identification result, and the number of product types and the number of products in the second identification result; By comparing the number of product types in the first identification result and the number of product types in the second identification result, the identification result with more product types is taken as the target identification result; If the number of product types in the first identification result is the same as the number of product types in the second identification result, then the identification result with the largest number of products in both the first and second identification results shall be taken as the target identification result.
7. The method according to claim 1, characterized in that, After analyzing the first identification result and the second identification result to determine the target identification result, the method further includes: Obtain the product type and corresponding product quantity from the target identification results; The settlement of goods shall be carried out according to the types and quantities of the goods.
8. A video recognition device, characterized in that, Applied to a distribution system, the device includes: The camera module is used to capture the user's shopping process using cameras installed on the shelves, thus obtaining shopping videos; The first determining module is used to determine the first identifying end from multiple identifying ends; The first identification module is used to perform initial identification of user shopping behavior in the shopping video through the first identification terminal to obtain a first identification result; The second determining module is used to determine a second identification terminal from multiple identification terminals if the first identification result meets preset conditions. The second identification terminal is different from the first identification terminal. The preset conditions include the first identification result indicating that no successful user shopping was detected, or, if the first identification result includes an identification time, the identification time is not within a preset time range. The identification time refers to the time taken by the first identification terminal to identify the user's shopping behavior in the shopping video. The second recognition module is used to perform secondary recognition on the user's shopping behavior in the shopping video through the second recognition terminal to obtain a second recognition result; The analysis module is used to analyze the first identification result and the second identification result to determine the target identification result.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the video recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the video recognition method according to any one of claims 1 to 7.