Warehouse scene recognition method and device, storage medium and electronic equipment
By combining image content understanding and element tracking and positioning recognition technologies, the problem of failure to recognize warehouse scenes without prior recording was solved, achieving accurate identification of warehouse scenes and location of abnormal situations, thus improving the efficiency of intelligent warehouse management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING GEEKPLUS TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot accurately identify unrecorded warehousing scenarios, leading to identification failures.
By acquiring image data of the warehouse site, the first model is used to understand and recognize the image content to obtain scene semantic information, and the second model is used to track and locate image elements to obtain element tracking information. The scene semantic information and element tracking information are then matched to obtain the warehouse scene recognition result.
It achieves accurate identification and location tracking of unrecorded warehousing scenarios, improves scenario understanding and fault identification generalization capabilities, enables timely handling of abnormal situations, and enhances the efficiency of intelligent warehouse management.
Smart Images

Figure CN122116359A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of intelligent warehousing technology, specifically to a method, apparatus, storage medium, and electronic device for identifying warehousing scenarios. Background Technology
[0002] During the operation of unmanned intelligent warehouses, it is necessary to understand the situation or state of various warehousing scenarios in order to better manage the intelligent warehouses.
[0003] Currently, related technologies can pre-record some warehousing scenarios. When warehousing scenario recognition is needed, the content of the occurring scenario is matched with these pre-recorded scenarios. If a match is successful, the recognition result is determined based on the matched scenario. However, if the content of the occurring scenario is not pre-recorded, the related technologies will be unable to accurately identify the warehousing scenario. Summary of the Invention
[0004] In view of this, the present disclosure provides a method, device, storage medium and electronic device for identifying warehousing scenarios, the main purpose of which is to improve the technical problem that if the content of the scenario is not recorded in advance, the relevant technology cannot accurately identify the warehousing scenario.
[0005] Firstly, this disclosure provides a method for identifying warehouse scenarios, including:
[0006] Acquire image data captured at the warehouse site;
[0007] The image data is input into a first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data; and,
[0008] The image data is input into the second model to track, locate, and identify image elements, thereby obtaining the element tracking information corresponding to the image data;
[0009] The warehouse scene identification result is obtained by matching and combining the scene semantic information and the element tracking information.
[0010] Optionally, the step of matching and combining the scene semantic information and the element tracking information to obtain the warehouse scene recognition result includes:
[0011] Obtain the event object from the scene semantic information;
[0012] Analyze the location information corresponding to the event object based on the element tracking information;
[0013] The location information corresponding to the event object is combined with the scene semantic information to obtain the warehouse scene recognition result.
[0014] Optionally, analyzing the location information corresponding to the event object based on the element tracking information includes:
[0015] Determine the target image element corresponding to the event object;
[0016] Obtain the coordinate information corresponding to the target image element from the element tracking information;
[0017] The location information corresponding to the event object is determined based on the coordinate information.
[0018] Optionally, combining the location information corresponding to the event object with the scene semantic information to obtain the warehouse scene recognition result includes:
[0019] From the scene semantic information, obtain the semantic information of the events that occur in the warehousing scene for the event object;
[0020] The semantic information of the event that occurs in the warehouse scenario is combined with the location information to obtain event information that includes the location information;
[0021] The warehouse scene identification result is determined based on the event information.
[0022] Optionally, the first model includes a large language model for warehouse image understanding, the training process of which includes:
[0023] Collect sample image data from the warehouse environment;
[0024] The sample image data is labeled, and the labeling includes at least one of element labeling, action labeling, relationship labeling, and event labeling;
[0025] A training set is constructed using labeled sample image data;
[0026] The large language model is obtained by training the training set.
[0027] Optionally, the step of inputting the image data into the first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data includes:
[0028] The image data is encoded into encoded data that the large language model can recognize;
[0029] The target data is input into the large language model for image content understanding and recognition to obtain the scene semantic information, wherein the target data includes one of the following:
[0030] The encoded data;
[0031] The encoded data and the user-inputted requirements.
[0032] Optionally, the step of inputting the image data into the second model for tracking, locating, and recognizing image elements to obtain element tracking information corresponding to the image data includes:
[0033] The second model is used to extract image elements present in the image data, determine the pixel information of the image elements in the image data, and perform positioning and tracking of the image elements based on the pixel information and the image frames of the image data to obtain the element tracking information of the image elements in the image data.
[0034] Optionally, the image data is fused from images captured in multiple warehouse scenarios, and the method further includes:
[0035] The image data is segmented into scenes to obtain image fragments of different warehouse scenes;
[0036] The step of inputting the image data into the first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data includes:
[0037] Input any image fragment of a target warehousing scene into the first model to understand and recognize the image content, and obtain the scene semantic information corresponding to the target warehousing scene;
[0038] The step of inputting the image data into the second model for tracking, locating, and recognizing image elements to obtain element tracking information corresponding to the image data includes:
[0039] The image fragments of the target warehousing scene are input into the second model for tracking, locating and recognizing image elements to obtain element tracking information corresponding to the target warehousing scene;
[0040] The process of matching and combining the scene semantic information and the element tracking information to obtain the warehouse scene recognition result includes:
[0041] The identification result of the target warehousing scene is obtained by matching and combining the scene semantic information and element tracking information corresponding to the target warehousing scene.
[0042] Secondly, this disclosure provides a method for identifying warehousing scenarios, including:
[0043] Receive warehouse scene recognition results, wherein the warehouse scene recognition results are obtained by matching and combining scene semantic information and element tracking information corresponding to image data captured on the warehouse site. The scene semantic information is obtained by image content understanding and recognition through a first model, and the element tracking information is obtained by image element tracking and positioning recognition through a second model.
[0044] Output is based on the warehouse scene recognition results.
[0045] Optionally, the first model includes a large language model for warehouse image understanding;
[0046] Before receiving the warehouse scene recognition result, the method further includes:
[0047] Obtain the user's input request information;
[0048] The requirement information is sent, and the requirement information is used to combine the image data with the large language model to understand and recognize the image content, thereby obtaining the scene semantic information.
[0049] Optionally, the output based on the warehouse scene recognition result includes:
[0050] Based on the warehouse scene recognition results, the target image corresponding to the event information in the image data is output. The event information is obtained by combining the semantic information of the event that occurred in the warehouse scene with the location information corresponding to the event object. The target image is marked with the location information corresponding to the event object.
[0051] Optionally, based on the warehouse scene recognition result, the target image corresponding to the event information in the image data is output, including:
[0052] If an abnormal situation is determined based on the warehouse scene recognition results, a target image marked with the location information of the abnormal event object is output.
[0053] Thirdly, this disclosure provides a warehouse scene identification device, including:
[0054] The acquisition module is configured to acquire image data captured at the warehouse site.
[0055] The recognition module is configured to input the image data into a first model to understand and recognize the image content, thereby obtaining scene semantic information corresponding to the image data; and to input the image data into a second model to track, locate, and recognize image elements, thereby obtaining element tracking information corresponding to the image data.
[0056] The generation module is configured to match and combine the scene semantic information and the element tracking information to obtain the warehouse scene recognition result.
[0057] Fourthly, this disclosure provides a recognition device for warehouse scenarios, including:
[0058] The transmission module is configured to receive the warehouse scene recognition result, wherein the warehouse scene recognition result is obtained by matching and combining the scene semantic information and element tracking information corresponding to the image data captured on the warehouse site. The scene semantic information is obtained by image content understanding and recognition through a first model, and the element tracking information is obtained by image element tracking and positioning recognition through a second model.
[0059] The output module is configured to output based on the warehouse scene recognition results.
[0060] Fifthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the warehouse scene identification method described in the first or second aspect.
[0061] In a sixth aspect, this disclosure provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the warehouse scene identification method described in the first or second aspect.
[0062] In a seventh aspect, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the warehouse scene identification method described in the first or second aspect.
[0063] By employing the above technical solution, this disclosure provides a method, apparatus, storage medium, and electronic device for identifying warehouse scenes. Compared with current related technologies, this disclosure's technical solution does not require prior recording of the warehouse scene. Specifically, image data captured at the warehouse site can be input into a first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data. The image data can also be input into a second model for image element tracking and location recognition to obtain element tracking information corresponding to the image data. Furthermore, the scene semantic information and element tracking information can be matched and combined to obtain the warehouse scene identification result. By applying this disclosure's technical solution, the occurrence of warehouse scenes can be accurately identified, and the location of the situation can be tracked and located, facilitating users to promptly understand the scene content and handle abnormal situations based on their location, thereby better managing intelligent warehouses.
[0064] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, specific embodiments of this disclosure are described below. Attached Figure Description
[0065] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0066] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0067] Figure 1 A flowchart illustrating a method for identifying a warehouse scenario provided in an embodiment of this disclosure is shown.
[0068] Figure 2 A flowchart illustrating a method for identifying a warehouse scenario provided in an embodiment of this disclosure is shown.
[0069] Figure 3 A flowchart illustrating a method for identifying a warehouse scenario provided in an embodiment of this disclosure is shown.
[0070] Figure 4 A flowchart illustrating a method for identifying a warehouse scenario provided in an embodiment of this disclosure is shown.
[0071] Figure 5 A flowchart illustrating a method for identifying a warehouse scenario provided in an embodiment of this disclosure is shown.
[0072] Figure 6 A schematic diagram of the structure of a warehouse scene identification device provided in an embodiment of this disclosure is shown;
[0073] Figure 7 A schematic diagram of the structure of a warehouse scene identification device provided in an embodiment of this disclosure is shown;
[0074] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0075] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0076] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, this disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this disclosure. Therefore, this disclosure is not limited to the specific implementations disclosed below.
[0077] The terminology used in one or more embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this disclosure. The singular forms “a,” “the,” and “the” as used in one or more embodiments of this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0078] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when…”.
[0079] To address the technical problem that current technologies cannot accurately identify warehouse scenarios if the content of the scenario is not recorded in advance, this disclosure provides a method for identifying warehouse scenarios, which can be applied to server-side execution for warehouse scenario identification, such as... Figure 1 As shown, the method includes the following steps 101 to 103:
[0080] Step 101: Obtain image data captured at the warehouse site.
[0081] In some embodiments, image data can be obtained by capturing images of the warehouse site using cameras installed in the warehouse environment. This image data may include images of the appearance of goods, the quantity of goods, the packaging of goods, the distribution of warehouse locations, images of shelves, images of cargo boxes, images of robots in operation, images of forklifts in operation, images of workstations in operation, images of conveyor belts in operation, images of personnel entering and leaving the warehouse, images of personnel working, and monitoring images of main aisles and safety passages within the warehouse, etc.
[0082] In some examples, the image data captured at the warehouse site can be video data or image data, etc.
[0083] In some examples, the corresponding shooting time and date can be recorded along with the image data, which helps to trace the specific time of the event when anomalies are found through warehouse scene identification results.
[0084] In some examples, image data can be sent to a warehouse scene recognition server to execute the recognition method shown in steps 101 to 103.
[0085] Step 102: Input the image data into the first model to understand and recognize the image content, and obtain the scene semantic information corresponding to the image data; input the image data into the second model to track, locate and recognize the image elements, and obtain the element tracking information corresponding to the image data.
[0086] In some embodiments, the first model can be used to understand and recognize image content to obtain scene semantic information. For example, image features are extracted from image data, and objects in the image are identified and classified based on the extracted features, such as identifying information like boxes, shelves, and robots in the image. Based on this image recognition, the structure and layout of the warehouse environment can be further analyzed, such as identifying spatial information like the arrangement of shelves, the way goods are placed, and the direction of aisles. Scene semantic information, such as changes in the state of goods, the handling of boxes, the operation of equipment, and personnel operations, is extracted from the analyzed warehouse environment. The extracted scene semantic information can be used to determine the situation occurring at the warehouse site.
[0087] In some embodiments, the second model can be used to track, locate, and identify image elements (such as boxes, shelves, goods, equipment, charging stations, workstations, pallets, cells, areas, etc.) in an image to obtain element tracking information. For example, it tracks each element appearing in the image data, determines the position of these elements, and identifies what the elements are. These elements are continuously tracked in consecutive images, such as tracking the position and size of elements in each frame, and handling potential issues like occlusion and deformation. Then, methods such as smoothing filtering and motion model prediction are used to optimize the tracking results, improving the stability and accuracy of the tracking.
[0088] In some embodiments, the first model and the second model may be two modules included in a preset model. The preset model integrates the first model and the second model. The preset model may be a multimodal large model. The preset model can be used to identify the corresponding warehouse scene recognition result based on image data captured on the warehouse site. When the image data is input into the preset model, it can be input into the first model and the second model respectively.
[0089] It should be noted that the names of the first model and the second model are not limited in this embodiment. For example, the first model may be called a warehouse image understanding model or an image understanding model, and the second model may be called a positioning and tracking model, a tracking model, or a positioning model.
[0090] Step 103: Match and combine the scene semantic information and element tracking information corresponding to the image data to obtain the warehouse scene recognition result.
[0091] In this embodiment, image understanding (scene semantic information) and tracking and positioning (element tracking information) of image data can be combined. The semantic information of an element in the image understanding is supplemented with the corresponding positioning and tracking information of that element. Then, the content of the warehouse scene is determined using the semantic information with positioning and tracking information, such as identifying the situation and tracing its location. For example, the scene semantic information may include event content related to a certain element A in the image data, such as a cargo box falling from a height. Adding the corresponding positioning and tracking information of element A allows for further determination of the location where the cargo box fell, given that it was identified as having fallen from a height.
[0092] As an alternative approach, after the first model outputs scene semantic information for image understanding and the second model outputs element tracking information for image data, the relevant outputs of the two models can be matched and combined to obtain warehouse scene semantics containing coordinate information, such as determining where situation A occurred.
[0093] For example, taking video data as an example, the target element is first extracted from the video understanding output by the first model. When the first model outputs the semantics of the scene video data, such as the semantics of "a box falling on the ground," the image elements obtained from this semantics are "box" and "ground." Then, based on the element tracking information output by the second model, the location information of the target element is determined. For example, based on the localization and tracking of image elements, other elements that are not related to "box" and "ground," such as robots and shelves, can be removed, leaving only the tracking information of "box" and "ground." Then, the location information of the target element is combined with the semantic understanding to obtain the warehouse scene information containing location information. For example, combining the tracking information of "box" and "ground" with the semantic understanding yields the warehouse scene information of "the box fell on the ground at coordinate ××" or "the box fell from position ×× onto the ground at coordinate ××."
[0094] Compared to current related technologies, the embodiments of this disclosure do not require prior recording of the warehouse scene. Specifically, image data captured at the warehouse site can be input into a first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data. The image data can also be input into a second model for image element tracking and location recognition to obtain element tracking information. Furthermore, the scene semantic information and element tracking information can be matched and combined to obtain the warehouse scene recognition result. By applying the technical solution of the embodiments of this disclosure, the occurring warehouse scene can be accurately identified, and the location of the situation can be tracked and located. This allows users to promptly understand the scene content and handle abnormal situations based on their location, thereby improving the management of intelligent warehousing.
[0095] Furthermore, as an alternative to step 103, such as Figure 2 As shown, step 103 may specifically include the methods shown in steps 201 to 203 below:
[0096] Step 201: Obtain the event object from the scene semantic information.
[0097] Scene semantic information can include a semantic understanding of the content of warehouse images, describing what happened on-site, such as changes in the state of goods and the operation of equipment. Based on this semantic understanding, it's necessary to define the types of events to be identified. For example, in warehouse videos, events of interest might include goods entering and leaving the warehouse, handling, and inventory checks. Further analysis of the scene semantics is needed to identify events matching predefined event types. If an event is identified, objects related to the event need to be extracted from the scene semantics. These objects could be participants in the event (e.g., people, goods), tools (e.g., forklifts, barcode scanners), or the environment (e.g., shelves, warehouse areas). By analyzing the event's behavior in the image (e.g., actions, changes in position), objects related to the event, i.e., event objects, can be inferred.
[0098] Step 202: Analyze the location information corresponding to the event object based on the element tracking information.
[0099] In some embodiments, step 202 may specifically include: first determining the target image element corresponding to the event object, such as determining the target image element corresponding to the event object in the image data; then obtaining the coordinate information corresponding to the target image element in the element tracking information (obtained by the second model); and then determining the position information corresponding to the event object based on the coordinate information corresponding to the target image element.
[0100] For example, if the event object is cargo A, determine the image element a corresponding to cargo A in the image data, then obtain the coordinate information corresponding to the image element a from the element tracking information, and then determine the location information corresponding to cargo A based on the coordinate information corresponding to the image element a.
[0101] Step 203: Combine the location information corresponding to the event object with the scene semantic information to obtain the warehouse scene recognition result.
[0102] In some embodiments, step 203 may specifically include: first, obtaining semantic information of the event that occurred in the warehouse scene from the scene semantic information; then combining the semantic information of the event that occurred in the warehouse scene with the location information (the location information obtained in step 202) to obtain event information containing the location information; and then determining the warehouse scene recognition result based on the event information.
[0103] For example, semantic information about events that occur to goods A in a warehousing scenario can be obtained from scene semantic information, such as semantic information that goods A has no outer packaging. This semantic information can then be combined with the location information corresponding to goods A to obtain event information that goods A has no outer packaging at location ××. Based on this event information, the scene recognition result of the current warehousing site can be determined.
[0104] Compared with current related technologies, the embodiments of this disclosure can combine scene semantic information and element tracking information to obtain warehouse scene recognition results, thereby accurately identifying the warehouse scene in question. This improves scene understanding and the generalization ability of fault identification, eliminating the need to pre-enter scene information into the model. It can also recognize new scene types, increasing the scene recognition rate. Furthermore, it can track and locate the position of the occurrence, i.e., based on the extraction and tracking of image elements, it can obtain the specific coordinates of the scene occurrence, allowing users to understand the scene content in a timely manner and handle the abnormal situation according to its location. To a certain extent, this can greatly improve intelligent warehouse management and maintenance.
[0105] Furthermore, in some examples, the first model may include a large language model (LLM) for warehouse image understanding, into which image data is input for understanding and recognizing image content in order to obtain scene semantic information corresponding to the image data.
[0106] In some examples, the training process of a large language model for warehouse image understanding includes: first, collecting sample image data from the warehouse environment; then, labeling these sample image data, where the labeling may include at least one of element labeling, action labeling, relation labeling, and event labeling; then, constructing a training set using the labeled sample image data; and finally, training the aforementioned large language model based on the training set. This approach yields a large language model that meets the requirements of the scenario, facilitating accurate warehouse scene recognition in the future.
[0107] Furthermore, in step 102, the image data is input into the first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data. Specifically, this may include the processes shown in steps 301 to 302 below, such as... Figure 3 As shown:
[0108] Step 301: Encode the image data into encoded data that can be recognized by the large language model for warehouse image understanding.
[0109] For example, image data is encoded using a video encoder (CLIP), and the encoded data is further processed by a multilayer perceptron to obtain image encoding, which is encoded data that can be recognized by a large language model for warehouse image understanding.
[0110] Step 302: Input the target data into the large language model for warehouse image understanding to understand and recognize the image content and obtain scene semantic information.
[0111] In some embodiments, the target data includes one of the following a to b:
[0112] a. Encoded data (obtained in step 301). For example, image data is encoded into data that can be recognized by a large language model, and then the encoded data is input into the large language model for semantic understanding to obtain the scene semantics corresponding to the image data, such as outputting what happened at the warehouse site.
[0113] b. Encoded data (obtained in step 301) and user-inputted requirements. For example, image data is encoded into data that a large language model can recognize. Then, the encoded data, along with the user's requirements in the intelligent dialogue system, is input into the large language model for semantic understanding to obtain the scene semantics corresponding to the image data, and output accordingly based on the user's requirements.
[0114] In some embodiments, step 102, which involves inputting image data into a second model for tracking, locating, and recognizing image elements to obtain element tracking information corresponding to the image data, may specifically include: extracting image elements present in the image data through the second model, determining the pixel information of the image elements in the image data, and locating and tracking the image elements based on the pixel information and the image frames of the image data to obtain element tracking information of the image elements in the image data.
[0115] For example, in the extraction and tracking of image elements, image data is classified and extracted to determine the semantics of the elements, and tracking is performed based on pixels to obtain the movement trajectory and coordinate information of the image elements. For instance, extracting and tracking image elements contained in scene image data yields the coordinate information of the image elements contained in the scene video data. Taking video data as an example, extracting and tracking image elements contained in the scene video data to obtain their coordinate information includes: extracting image elements from the scene video data based on a classification network; performing semantic understanding and matching of image elements based on a semantic matching network; determining the pixel information of image elements based on a region optimization network; and tracking the image elements based on the pixel information and the image frames of the scene video data using a target tracking network to obtain the coordinate information of the image elements contained in the scene video data. For example, the video image includes elements such as boxes, shelves, robots, and the ground; after this step, the coordinate information of elements such as boxes, shelves, robots, and the ground will be obtained. If any element moves, its movement trajectory can also be obtained.
[0116] Furthermore, by obtaining scene semantic information and element tracking information through the above methods, the scene semantic information and element tracking information are matched and combined based on the matching network to obtain the warehouse scene recognition result.
[0117] By applying the technical solutions of the above embodiments, scene semantic information and element tracking information corresponding to image data can be accurately obtained. Then, based on the scene semantic information and element tracking information, matching and combination can be performed to accurately identify the warehouse scene that occurred, and the location of the situation can be tracked and located. This makes it easier for users to understand the scene content in a timely manner and to handle the abnormal situation according to its location, thereby better managing the intelligent warehouse.
[0118] Furthermore, if the image data to be identified is video data, the acquired source video data can be segmented according to different scenes, so that a whole segment of source video data is divided into video segments of different scenes, and each video segment is a scene video data. Specifically, the process shown in steps 401 to 404 can be executed as follows: Figure 4 As shown:
[0119] Step 401: Obtain image data captured at the warehouse site. This image data is a fusion of images captured in multiple warehouse scenarios.
[0120] Step 402: Segment the image data to obtain image fragments of different warehousing scenarios.
[0121] To facilitate the identification of semantic information in each scene and the tracking of elements within that scene, this embodiment of the present disclosure inputs video segment data of a single scene into the first and second models respectively. Therefore, after obtaining the large original source video data of the intelligent warehouse (including video segment data from multiple scenes), scene segmentation can be performed first based on a scene segmentation network to divide the source video data containing video segment data from multiple scenes into multiple video segment data of different scenes.
[0122] After obtaining the video segment data for each scene, the video segment data for the same scene is input into the first model and the second model respectively, i.e., the process shown in step 403 is executed.
[0123] Step 403: Input any image fragment of the target warehousing scene into the first model to understand and recognize the image content, and obtain the scene semantic information corresponding to the target warehousing scene; input the image fragment of the target warehousing scene into the second model to track and locate the image elements, and obtain the element tracking information corresponding to the target warehousing scene.
[0124] Step 404: Match and combine the scene semantic information and element tracking information corresponding to the target warehousing scene to obtain the recognition result of the target warehousing scene.
[0125] This embodiment combines video understanding and tracking / localization when processing video data to obtain the coordinate information of elements within the video semantic understanding. This facilitates subsequent intelligent warehouse management, such as fault recovery. The first model performs semantic understanding of the video data; the video data is input into the first model to obtain the semantic meaning expressed by the corresponding video data, such as a cargo box falling to the ground. The second model extracts and tracks video image elements, classifying and extracting elements from the video data, determining the semantic meaning of each element, and tracking them based on pixels to obtain the movement trajectory and coordinate information of the video image elements. After the first model outputs the understood scene video semantics and the element tracking information of the scene video data, the relevant output results can be matched and combined to obtain semantics containing coordinate information, i.e., determining where something happened and what happened. This allows for accurate identification of the warehouse scene, improving scene understanding and the generalization ability of fault identification. It eliminates the need to pre-enter scene information into the model and can also identify new scene types, thus improving the scene recognition rate. Furthermore, it can track and locate the location of the situation, that is, based on the extraction and tracking of image elements, it can obtain the specific coordinates of the scene, which makes it easier for users to understand the scene content in a timely manner and to handle the abnormal situation according to its location, which can greatly improve intelligent warehouse management and maintenance to a certain extent.
[0126] Furthermore, after obtaining the warehouse scene recognition result through the above embodiments, the server can send the warehouse scene recognition result to the client so that the warehouse scene recognition result can be output and displayed. Accordingly, to illustrate the specific execution process on the client side, Figure 5 A schematic diagram of a method for identifying a warehouse scenario according to an embodiment of the present disclosure is shown. The method is executed on the client side and may include the steps shown in 501 to 502:
[0127] Step 501: Receive the warehouse scene recognition results.
[0128] In some embodiments, the client can receive the warehouse scene recognition results sent by the server. The server can either serve as the client's backend service support function and act as the client's backend service server, or it can be an independent server that provides various online services and provides online service resources for numerous clients or other related applications.
[0129] The warehouse scene recognition result is obtained by matching and combining scene semantic information and element tracking information corresponding to image data captured at the warehouse site. The scene semantic information is obtained through image content understanding and recognition using a first model, while the element tracking information is obtained through image element tracking and positioning recognition using a second model. For details of the recognition process, please refer to [link to relevant documentation]. Figures 1 to 4 The description of the example shown will not be repeated here.
[0130] Step 502: Output based on the warehouse scene recognition results.
[0131] In some embodiments, after obtaining the warehouse scene recognition results, the results can be output, such as through images, text, video, audio, light, vibration, etc., to output the recognized warehouse scene information, so that users can intuitively understand the current situation of the warehouse scene.
[0132] In some examples, the output of warehouse scene information can be proactive or user-triggered. For instance, when an abnormal situation is detected in a warehouse scene, only an error message is displayed. The user can then take further action or make requests based on this error message before the specific content of the abnormal scene is output. This could be an image of the abnormal scene, text about the abnormal scene, or a video of the abnormal scene.
[0133] Compared to current related technologies, the embodiments of this disclosure do not require prior recording of the warehouse scene. Specifically, image data captured at the warehouse site can be input into a first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data. The image data can also be input into a second model for image element tracking and location recognition to obtain element tracking information. Furthermore, the scene semantic information and element tracking information can be matched and combined to obtain the warehouse scene recognition result. By applying the technical solution of the embodiments of this disclosure, the occurring warehouse scene can be accurately identified, and the location of the situation can be tracked and located. This allows users to promptly understand the scene content and handle abnormal situations based on their location, thereby improving the management of intelligent warehousing.
[0134] Furthermore, in some embodiments, the first model may include a large language model for warehouse image understanding. Before step 501, user-inputted demand information may be obtained and sent to the server. This demand information is used to combine image data with the large language model to understand and recognize image content, thereby obtaining scene semantic information.
[0135] In some examples, an intelligent dialogue system can be built based on a large language model. Users can use this intelligent dialogue system on the client side to make requests and input relevant request information, such as asking whether there are any anomalies in the current warehouse scene, what kind of anomalies have occurred, and the required data format. Image data is also input into the large language model for image content understanding and recognition to obtain the scene semantic information corresponding to the image data. Furthermore, the element tracking information output by the second model can be combined to obtain the warehouse scene content containing anomaly location information.
[0136] For example, in a paper cup picking scenario, there are multiple paper cups in the cargo box. The corresponding captured image data can be input into a large language model, and the dialogue between the user and the system would be as follows:
[0137] User inquiry: (What's inside the cargo box? Output format: Category: Quantity);
[0138] System response: (Category: Paper cup, Quantity: 4);
[0139] User inquiry: (Are there any damaged items? Output format: Category: Yes / No);
[0140] System response: (Paper cup: Yes);
[0141] User inquiry: (What type of damage is it?)
[0142] The system responded: (There is a bunch of cups with broken plastic packaging, the packaging is torn, and the cups are exposed), and can also attach an image showing the location of the broken plastic packaging;
[0143] User request: (Select an item to pick up);
[0144] The system replied: (Since there is a damaged plastic package, it is best to choose the other pile with intact packages), and you can also attach an image showing the location of the intact plastic packages.
[0145] For example, in a surveillance scenario, the boxes on the shelves may contain items such as clothing. The captured image data can be input into a large language model, and the dialogue between the user and the system would be as follows:
[0146] User question: (You are responsible for warehouse scene monitoring. Please describe any unusual events or objects in the warehouse.)
[0147] System response: (The robot has stopped);
[0148] User asked: (Why did the robot stop?)
[0149] System response: (Description: A garment in its packaging is falling off the shelf. Object: A garment in its packaging.)
[0150] Behavior: Falling from the shelf.
[0151] User inquiry: (Location: Clothes with packaging);
[0152] The system responded: Displaying an image showing the exact location of the garment.
[0153] For example, in a surveillance scenario, the captured surveillance image data can be input into a large language model, and the user's interaction with the system can be as follows:
[0154] User question: (You are a warehouse scene monitor, and you need to calculate the number of people entering the warehouse and the number of people staying in the warehouse. How many people entered the warehouse in the past hour? Output the data in the format (Total number of people: Quantity, Current number of people: Quantity)).
[0155] System response: (Total number of users: 2, Current number of users: 1.)
[0156] User inquiry: (Locate the person who left.)
[0157] The system responded: It displays the surveillance footage of the person leaving and marks the person who left.
[0158] Through the above embodiments, users can easily and quickly locate the situation through the intelligent dialogue system, intuitively understand the scene, and facilitate effective management.
[0159] In some embodiments, after obtaining the warehouse scene recognition result, in order to enable users to intuitively understand the location of the target object in the situation, the target image corresponding to the event information in the image data can be output based on the warehouse scene recognition result. The target image can be a picture or a video. The event information is obtained by combining the semantic information of the event that occurred in the warehouse scene with the location information corresponding to the event object. For details, please refer to the description of the example shown in step 203, which will not be repeated here. The target image is marked with the location information corresponding to the event object.
[0160] For example, the event information could be that goods A at location ×× do not have outer packaging. Based on the captured image data, an image of goods A can be output, and goods A can be labeled in the image so that users can intuitively understand that goods A does not have outer packaging.
[0161] Furthermore, in some examples, based on the warehouse scene recognition results, the target image corresponding to the event information in the image data is output. Specifically, this may include: if the warehouse scene recognition results determine that an abnormal situation has occurred, outputting a target image marked with the location information of the abnormal event object. For example, if the warehouse scene recognition results determine that a packaged item has fallen from the shelf, an image marked with the specific location of the item can be output.
[0162] In this way, users can intuitively understand the abnormal situation and quickly locate the location of the abnormal situation, which facilitates timely handling and achieves effective management.
[0163] Furthermore, as Figures 1 to 4 The specific implementation of the example shown in this disclosure provides a warehouse scene recognition device that can be applied to the server side, such as... Figure 6 As shown, the device includes: an acquisition module 61, an identification module 62, and a generation module 63.
[0164] The acquisition module 61 is configured to acquire image data captured at the warehouse site.
[0165] The recognition module 62 is configured to input the image data into a first model to understand and recognize the image content, thereby obtaining scene semantic information corresponding to the image data; and to input the image data into a second model to track, locate, and recognize image elements, thereby obtaining element tracking information corresponding to the image data.
[0166] The generation module 63 is configured to match and combine the scene semantic information and the element tracking information to obtain the warehouse scene recognition result.
[0167] In some embodiments, the generation module 63 is specifically configured to: obtain event objects in the scene semantic information; analyze the location information corresponding to the event objects based on the element tracking information; and combine the location information corresponding to the event objects with the scene semantic information to obtain the warehouse scene recognition result.
[0168] In some embodiments, the generation module 63 is further configured to: determine the target image element corresponding to the event object; obtain coordinate information corresponding to the target image element from the element tracking information; and determine the location information corresponding to the event object based on the coordinate information.
[0169] In some embodiments, the generation module 63 is further configured to: obtain semantic information of the events that occur in the warehousing scene from the scene semantic information; combine the semantic information of the events that occur in the warehousing scene with the location information to obtain event information containing the location information; and determine the warehousing scene recognition result based on the event information.
[0170] In some embodiments, the first model includes a large language model for warehouse image understanding; the acquisition module 61 is further configured to collect sample image data in the warehouse environment; to annotate the sample image data, the annotation including at least one of element annotation, action annotation, relation annotation, and event annotation; to construct a training set using the annotated sample image data; and to train the large language model based on the training set.
[0171] In some embodiments, the recognition module 62 is specifically configured to encode the image data into encoded data that the large language model can recognize; input the target data into the large language model to understand and recognize the image content, thereby obtaining the scene semantic information, wherein the target data includes one of the following:
[0172] The encoded data; the encoded data and the user-inputted requirement information.
[0173] In some embodiments, the identification module 62 is further configured to extract image elements present in the image data through the second model, determine the pixel information of the image elements in the image data, and perform localization and tracking of the image elements based on the pixel information and the image frames of the image data to obtain element tracking information of the image elements in the image data.
[0174] In some embodiments, the image data is fused from images captured in multiple warehousing scenarios, and the acquisition module 61 is further configured to perform scene segmentation on the image data to obtain image fragments of different warehousing scenarios;
[0175] The recognition module 62 is further configured to input an image fragment of any target warehousing scene into the first model for understanding and recognizing the image content to obtain scene semantic information corresponding to the target warehousing scene; and to input an image fragment of the target warehousing scene into the second model for tracking, locating and recognizing image elements to obtain element tracking information corresponding to the target warehousing scene.
[0176] The generation module 63 is further configured to match and combine the scene semantic information and element tracking information corresponding to the target warehousing scene to obtain the recognition result of the target warehousing scene.
[0177] It should be noted that other corresponding descriptions of the functional units involved in the warehouse scene identification device provided in this disclosure embodiment can be found in the following references. Figures 1 to 4 The corresponding descriptions in [the document] will not be repeated here.
[0178] Furthermore, as Figure 5 The specific implementation of the example shown in this disclosure provides a warehouse scene recognition device that can be applied to a client, such as... Figure 7 As shown, the device includes: a transmission module 71 and an output module 72.
[0179] The transmission module 71 is configured to receive the warehouse scene recognition result, wherein the warehouse scene recognition result is obtained by matching and combining the scene semantic information and element tracking information corresponding to the image data captured at the warehouse site. The scene semantic information is obtained by image content understanding and recognition through a first model, and the element tracking information is obtained by image element tracking and positioning recognition through a second model.
[0180] The output module 72 is configured to output based on the warehouse scene recognition results.
[0181] In some embodiments, the first model includes a large language model for warehouse image understanding;
[0182] The transmission module 71 is also configured to acquire user-inputted demand information; send the demand information, which is used to combine the image data with the large language model to understand and recognize the image content, thereby obtaining the scene semantic information.
[0183] In some embodiments, the output module 72 is specifically configured to output a target image corresponding to the event information in the image data based on the warehouse scene recognition result, wherein the event information is obtained by combining the semantic information of the event that occurred in the warehouse scene with the location information corresponding to the event object, and the target image is marked with the location information corresponding to the event object.
[0184] In some embodiments, the output module 72 is further configured to output a target image marked with the location information of the abnormal event object when an abnormal situation is determined based on the warehouse scene recognition result.
[0185] It should be noted that other corresponding descriptions of the functional units involved in the warehouse scene identification device provided in this disclosure embodiment can be found in the following references. Figure 5 The corresponding descriptions in [the document] will not be repeated here.
[0186] Based on the above, Figures 1 to 4 As illustrated in the example, correspondingly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the above-described... Figures 1 to 4 The example method shown above. Based on the above... Figure 5 As illustrated in the example, correspondingly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the above-described... Figure 5 The example method shown.
[0187] Based on the above, Figures 1 to 4As illustrated, correspondingly, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described... Figures 1 to 4 The example method shown above. Based on the above... Figure 5 As illustrated, correspondingly, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described... Figure 5 The example method shown.
[0188] Based on this understanding, the technical solution of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive) and includes several instructions to cause a computer device (such as a personal computer, server, or network device) to execute the methods of various implementation scenarios of this disclosure.
[0189] According to embodiments of this disclosure, an electronic device is also provided.
[0190] Figure 8 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.
[0191] Electronic devices can also refer to various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0192] like Figure 8 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in ROM (Read-Only Memory) 1002 or loaded from storage unit 1008 into RAM (Random Access Memory) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. I / O (Input / Output) interface 1005 is also connected to bus 1004.
[0193] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0194] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as the identification methods for warehouse scenarios.
[0195] For example, in some embodiments, the warehouse scene identification method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the method described above can be performed. Alternatively, in other embodiments, computing unit 1001 can be configured to perform the aforementioned warehouse scene identification method by any other suitable means (e.g., by means of firmware).
[0196] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0197] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0198] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0199] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0200] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.
[0201] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0202] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0203] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0204] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for identifying warehouse scenarios, characterized in that, include: Acquire image data captured at the warehouse site; The image data is input into the first model to understand and recognize the image content, thereby obtaining the scene semantic information corresponding to the image data; and, The image data is input into the second model to track, locate, and identify image elements, thereby obtaining the element tracking information corresponding to the image data; The warehouse scene identification result is obtained by matching and combining the scene semantic information and the element tracking information.
2. The method according to claim 1, characterized in that, The process of matching and combining the scene semantic information and the element tracking information to obtain the warehouse scene recognition result includes: Obtain the event object from the scene semantic information; Analyze the location information corresponding to the event object based on the element tracking information; The location information corresponding to the event object is combined with the scene semantic information to obtain the warehouse scene recognition result.
3. The method according to claim 2, characterized in that, The step of analyzing the location information corresponding to the event object based on the element tracking information includes: Determine the target image element corresponding to the event object; Obtain the coordinate information corresponding to the target image element from the element tracking information; The location information corresponding to the event object is determined based on the coordinate information.
4. The method according to claim 2, characterized in that, The step of combining the location information corresponding to the event object with the scene semantic information to obtain the warehouse scene recognition result includes: From the scene semantic information, obtain the semantic information of the events that occur in the warehousing scene for the event object; The semantic information of the event that occurs in the warehouse scenario is combined with the location information to obtain event information that includes the location information; The warehouse scene identification result is determined based on the event information.
5. The method according to claim 1, characterized in that, The first model includes a large language model for warehouse image understanding, and the training process of the large language model includes: Collect sample image data from the warehouse environment; The sample image data is labeled, and the labeling includes at least one of element labeling, action labeling, relationship labeling, and event labeling; A training set is constructed using labeled sample image data; The large language model is obtained by training the training set.
6. The method according to claim 5, characterized in that, The step of inputting the image data into the first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data includes: The image data is encoded into encoded data that the large language model can recognize; The target data is input into the large language model for image content understanding and recognition to obtain the scene semantic information, wherein the target data includes one of the following: The encoded data; The encoded data and the user-inputted requirements.
7. The method according to claim 1, characterized in that, The step of inputting the image data into the second model for tracking, locating, and recognizing image elements to obtain element tracking information corresponding to the image data includes: The second model is used to extract image elements present in the image data, determine the pixel information of the image elements in the image data, and perform positioning and tracking of the image elements based on the pixel information and the image frames of the image data to obtain the element tracking information of the image elements in the image data.
8. The method according to claim 1, characterized in that, The image data is fused from images captured in multiple warehouse scenarios, and the method further includes: The image data is segmented into scenes to obtain image fragments of different warehouse scenes; The step of inputting the image data into the first model for image content understanding and recognition to obtain scene semantic information corresponding to the image data includes: Input any image fragment of a target warehousing scene into the first model to understand and recognize the image content, and obtain the scene semantic information corresponding to the target warehousing scene; The step of inputting the image data into the second model for tracking, locating, and recognizing image elements to obtain element tracking information corresponding to the image data includes: The image fragments of the target warehousing scene are input into the second model for tracking, locating and recognizing image elements to obtain element tracking information corresponding to the target warehousing scene; The process of matching and combining the scene semantic information and the element tracking information to obtain the warehouse scene recognition result includes: The identification result of the target warehousing scene is obtained by matching and combining the scene semantic information and element tracking information corresponding to the target warehousing scene.
9. A method for identifying warehouse scenarios, characterized in that, include: Receive warehouse scene recognition results, wherein the warehouse scene recognition results are obtained by matching and combining scene semantic information and element tracking information corresponding to image data captured on the warehouse site. The scene semantic information is obtained by image content understanding and recognition through a first model, and the element tracking information is obtained by image element tracking and positioning recognition through a second model. Output is based on the warehouse scene recognition results.
10. The method according to claim 9, characterized in that, The first model includes a large language model for warehouse image understanding; Before receiving the warehouse scene recognition result, the method further includes: Obtain the user's input request information; The requirement information is sent, and the requirement information is used to combine the image data with the large language model to understand and recognize the image content, thereby obtaining the scene semantic information.
11. The method according to claim 10, characterized in that, The output is based on the warehouse scene recognition results, including: Based on the warehouse scene recognition results, the target image corresponding to the event information in the image data is output. The event information is obtained by combining the semantic information of the event that occurred in the warehouse scene with the location information corresponding to the event object. The target image is marked with the location information corresponding to the event object.
12. The method according to claim 11, characterized in that, Based on the warehouse scene recognition results, the target image corresponding to the event information in the image data is output, including: If an abnormal situation is determined based on the warehouse scene recognition results, a target image marked with the location information of the abnormal event object is output.
13. A warehouse scene recognition device, characterized in that, include: The acquisition module is configured to acquire image data captured at the warehouse site. The recognition module is configured to input the image data into a first model to understand and recognize the image content, and obtain the scene semantic information corresponding to the image data; Furthermore, the image data is input into the second model for tracking, locating, and recognizing image elements to obtain element tracking information corresponding to the image data; The generation module is configured to match and combine the scene semantic information and the element tracking information to obtain the warehouse scene recognition result.
14. A warehouse scene recognition device, characterized in that, include: The transmission module is configured to receive the warehouse scene recognition result, wherein the warehouse scene recognition result is obtained by matching and combining the scene semantic information and element tracking information corresponding to the image data captured on the warehouse site. The scene semantic information is obtained by image content understanding and recognition through a first model, and the element tracking information is obtained by image element tracking and positioning recognition through a second model. The output module is configured to output based on the warehouse scene recognition results.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 12.
16. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 12.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 12.