Packing detection method and electronic equipment

By using machine learning models to evaluate operators' outer packaging inspection actions, the problem of irregular outer packaging inspection in electronic product production was solved, the outer packaging inspection was made intelligent and accurate, and the workload of rework was reduced.

CN118062365BActive Publication Date: 2025-09-26HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211478671.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-09-26
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

During the production process of electronic products, there are still irregular manual operations in the outer packaging inspection process, which leads to poor product quality control, inability to detect problems in a timely manner, and increased rework workload.

Method used

A machine learning model is used to identify the operator's actions, and video streams are collected through cameras to evaluate whether the operator's inspection actions meet the preset standards. If they do not meet the standards, the operator will be prompted to re-inspect in a timely manner, thereby improving the intelligence of outer packaging inspection.

Benefits of technology

It has achieved the disassembly of complex inspection scenarios, simplified the algorithm complexity, ensured the accuracy of test results, reduced reliance on manual supervision, improved the intelligence level of outer packaging inspection, and reduced the workload of rework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118062365B_ABST
    Figure CN118062365B_ABST
Patent Text Reader

Abstract

The present application provides a packing detection method and electronic equipment, which relate to the field of artificial intelligence technology. Improve the intelligence level of the outer packaging inspection link. The specific scheme is: receive the first video stream captured by the camera; when the first video stream includes the first segment, evaluate whether the operator's first action meets the first standard based on the first segment and the preconfigured first model; the first segment records the process from the operator taking the first product into the first space to putting the first product into the first box. The first space is the acquisition space of the camera. The first box is used to place electronic products that have been inspected. The first box is located in the first space. The first action is the action of checking the outer packaging of the first product recorded in the first segment; when the first action does not meet the first standard, the first prompt information corresponding to the first identifier of the first product is displayed, prompting the operator to re-check the outer packaging of the first product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a packaging detection method and electronic equipment. Background Art

[0002] With the development of computer technology, people's lives have undergone tremendous changes. In particular, the application of artificial intelligence technology in industrial production has significantly improved the efficiency of industrial production.

[0003] Of course, in industrial production processes, such as the production of electronic products, there are still some steps that require manual work, such as the inspection of electronic product packaging before shipment. Manual packaging inspection is prone to errors and cannot detect problems in a timely manner, resulting in a large workload during rework and even affecting the quality control of electronic products. In other words, the packaging inspection process in the electronic product production process still lacks intelligent features. Summary of the Invention

[0004] In view of this, the present application provides a packing detection method and electronic equipment to improve the intelligence level of the outer packaging inspection link.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a packing detection method, which is applied to an electronic device, which is communicatively connected to a camera, and the method includes: receiving a first video stream captured by the camera; when the first video stream includes a first segment, evaluating whether the operator's first action meets the first standard based on the first segment and a preconfigured first model; wherein the first model is a machine learning model for identifying actions, and the first segment records the process from the operator taking the first product into the first space to putting the first product into the first box, the first space is the capture space of the camera, the first box is used to place electronic products that have been inspected, and the first box is located in the first space, and the first action is the action of checking the outer packaging of the first product recorded in the first segment; when the first action does not meet the first standard, a first prompt message corresponding to the first identifier of the first product is displayed, prompting the operator to recheck the outer packaging of the first product.

[0007] In the above embodiment, the operator inspects the outer packaging of the first product in the first space and places the first product in the first box after inspection. The camera can record the operator's actions from the time the first product is brought into the first space to the time it is placed in the first box, generating a first segment. In this way, the electronic device only needs to identify the first segment to analyze whether the operator effectively inspected the outer packaging of the first product before placing it in the first box. In other words, it can determine whether the first action in the first segment meets the first standard of the preset criteria. If the first standard is not met, the electronic device can promptly remind the operator to re-inspect the first product. This allows the operator to promptly identify any irregularities in the inspection of the first product and promptly rectify them, avoiding unnecessary rework and increased workload later. The entire process is completed by the electronic device, replacing manual supervision with artificial intelligence algorithms, improving the intelligence level of the outer packaging inspection process.

[0008] In some embodiments, after displaying the first prompt message, the method also includes: receiving a second video stream captured by the camera; when the second video stream includes a second segment, evaluating whether the operator's second action meets the first standard based on the second segment and the first model, wherein the second segment records the process of the operator taking the first product out of the first box and putting it back into the first box again, and the second action is the action of checking the outer packaging of the first product recorded in the second segment; if the second action meets the first standard, displaying a second prompt message to prompt the operator that the inspection of the first product is completed.

[0009] In the above embodiment, when the operator re-inspects the first product, they need to remove the first product from the first box and then return it to the first box after re-inspection. The camera can record the operator's movements from the time they remove the first product from the first box to the time they return it to the first box, generating a second segment. The electronic device then only needs to identify the second segment to assess whether the re-inspection action meets the first criterion.

[0010] Obviously, electronic equipment can disassemble complex inspection scenarios. Take the complex inspection scenario of taking the first product into the first space, putting the first product into the first box after inspection, and then taking the first product out of the first box for re-inspection as an example. The electronic equipment does not need to perform recognition learning for this type of scenario separately. It can directly split the complex inspection scenario into two inspection processes and perform identification separately. While ensuring accurate detection results, it simplifies the complexity of the algorithm.

[0011] In some embodiments, the method further includes: if the first action meets the first standard, displaying a second prompt message to prompt the operator that the inspection of the first product is completed.

[0012] In some embodiments, after displaying the second prompt message, the method also includes: determining a first quantity based on the video frames in the first video stream or the second video stream, the first quantity being the number of electronic products placed in the first box after the first product is placed in the first box; when the first quantity is equal to a preset first value and a third video stream captured by the camera is received, if the third video stream contains a third segment, then based on the third segment and the first model, evaluating whether the operator's third action meets the second standard, wherein the third segment records the operator's process of packing the first box, and the third action is the action of packing the first box recorded in the third segment; if the third action does not meet the second standard, displaying a third prompt message to prompt the operator to repack the first box.

[0013] In the above embodiment, through the recognition of the third segment, it is possible to analyze whether the operator's action of packing the box meets the preset standard, that is, the second standard. Similarly, if it is determined that the packing operation does not meet the second standard, the operator is promptly reminded to make corrections, thereby improving the packaging quality and improving the intelligence level of the outer packaging inspection link.

[0014] In some embodiments, after determining the first quantity, the method may further include: when the first quantity is equal to the first value, determining that the first box contains a first drawer strap based on video frames in the first video stream or the second video stream.

[0015] In some embodiments, after displaying the second prompt message, the method further includes: determining a first quantity based on the video frames in the first video stream or the second video stream, the first quantity being the number of electronic products placed in the first box after the first product is placed in the first box; when the first quantity is less than a preset first value and a third video stream captured by the camera is received, if the third video stream contains a third segment, displaying a fourth prompt message to indicate that the first box is not yet full, wherein the third segment records the process of the operator packing the first box.

[0016] In some embodiments, based on the first segment and the preconfigured first model, evaluating whether the operator's first action meets the first standard includes: inputting the first segment into the first model to obtain a first output result; and determining whether the first action meets the first standard based on the value of the first output result.

[0017] In some embodiments, before evaluating whether the operator's first action meets the first criterion based on the first segment and the preconfigured first model, the method also includes: when the first video stream includes a first video frame, starting from the first video frame, tracking the position of the first product in the first video stream until a second video frame is detected, wherein the first video frame is a video frame in which the first product appears in the first frame, and the first product in the second video frame is a video frame in which the first box is located; and determining that the video frames between the first video frame and the second video frame constitute the first segment.

[0018] In the above embodiment, the electronic device identifies the first segment by tracking the position change of the first product in the video stream, which is simple to implement and highly accurate.

[0019] In some embodiments, before evaluating whether the operator's second action meets the first criterion based on the second clip and the first model, the method also includes: when the second video stream includes a third video frame, starting from the third video frame, tracking the position of the first product in the second video stream until a fourth video frame is detected, wherein the center position of the first product in the third video frame does not overlap with the first box, the fourth video frame is a video frame in which the center position of the first product overlaps with the first box, and the acquisition time of the third video frame is earlier than that of the fourth video frame.

[0020] In the above embodiment, the electronic device identifies the second segment by tracking the position change of the first product in the video stream. This is simple to implement and highly accurate. In addition, it will be easier to add recognition of new types of video segments in the future.

[0021] In a second aspect, an embodiment of the present application provides an electronic device, which includes one or more processors and a memory; the memory is coupled to the processor, and the memory is used to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the one or more processors are used to: receive a first video stream captured by the camera; when the first video stream includes a first segment, evaluate whether the operator's first action meets the first standard based on the first segment and a preconfigured first model; wherein the first model is a machine learning model for identifying actions, and the first segment records the process from the operator taking the first product into the first space to putting the first product into the first box, the first space is the capture space of the camera, the first box is used to place electronic products that have been inspected, and the first box is located in the first space, and the first action is the action of checking the outer packaging of the first product recorded in the first segment; if the first action does not meet the first standard, a first prompt message corresponding to the first identifier of the first product is displayed, prompting the operator to recheck the outer packaging of the first product.

[0022] In some embodiments, after displaying the first prompt message, the one or more processors are used to: receive a second video stream captured by the camera; when the second video stream includes a second segment, evaluate whether the operator's second action meets the first standard based on the second segment and the first model, wherein the second segment records the process of the operator taking the first product out of the first box and putting it back into the first box again, and the second action is the action of checking the outer packaging of the first product recorded in the second segment; if the second action meets the first standard, display a second prompt message to prompt the operator that the inspection of the first product is completed.

[0023] In some embodiments, the one or more processors are configured to: when the first action meets the first standard, display a second prompt message to prompt the operator that the inspection of the first product is completed.

[0024] In some embodiments, after displaying the second prompt information, the one or more processors are used to: determine a first quantity based on the video frames in the first video stream or the second video stream, where the first quantity is the number of the electronic products placed in the first box after the first product is placed in the first box; when the first quantity is equal to a preset first value and a third video stream captured by the camera is received, if the third video stream contains a third segment, then based on the third segment and the first model, evaluate whether the operator's third action meets the second standard, wherein the third segment records the operator's process of packing the first box, and the third action is the action of packing the first box recorded in the third segment; if the third action does not meet the second standard, display a third prompt information to prompt the operator to repack the first box.

[0025] In some embodiments, after determining the first quantity, the one or more processors are used to: when the first quantity is equal to the first value, determine that the first box contains a first drawer strap based on the video frames in the first video stream or the second video stream.

[0026] In some embodiments, after displaying the second prompt information, the one or more processors are used to: determine a first quantity based on the video frames in the first video stream or the second video stream, where the first quantity is the number of electronic products placed in the first box after the first product is placed in the first box; when the first quantity is less than a preset first value and a third video stream captured by the camera is received, if the third video stream contains a third segment, display the fourth prompt information to indicate that the first box is not yet full, wherein the third segment records the process of the operator packing the first box.

[0027] In some embodiments, the one or more processors are configured to: input the first fragment into the first model to obtain a first output result; and determine whether the first action meets the first standard based on a value of the first output result.

[0028] In some embodiments, before evaluating whether the operator's first action meets the first criterion based on the first fragment and the preconfigured first model, the one or more processors are used to: when the first video stream includes a first video frame, start from the first video frame, track the position of the first product in the first video stream until a second video frame is detected, wherein the first video frame is the first frame in which the first product appears, and the second video frame is a frame in which the first product is located in the first box; and determine that the video frames between the first video frame and the second video frame constitute the first fragment.

[0029] In some embodiments, before evaluating whether the operator's second action meets the first criterion based on the second clip and the first model, the one or more processors are used to: when the second video stream includes a third video frame, track the position of the first product in the second video stream starting from the third video frame until a fourth video frame is detected, wherein the center position of the first product in the third video frame does not overlap with the first box, the fourth video frame is a video frame in which the center position of the first product overlaps with the first box, and the acquisition time of the third video frame is earlier than that of the fourth video frame.

[0030] In a third aspect, an embodiment of the present application provides a computer storage medium, comprising computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method in the above-mentioned first aspect and its possible embodiments.

[0031] In a fourth aspect, the present application provides a computer program product. When the computer program product is run on the above-mentioned electronic device, the electronic device executes the method in the above-mentioned first aspect and its possible embodiments.

[0032] It can be understood that the electronic devices, computer storage media and computer program products provided in the above aspects are all applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic diagram of the structure of a packing detection system provided in an embodiment of the present application;

[0034] Figure 2 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0035] Figure 3 This is one of the principle example diagrams of splitting video segment a provided in an embodiment of the present application;

[0036] Figure 4 The second diagram of the principle of dividing the video segment a provided in the embodiment of the present application;

[0037] Figure 5 One of the example diagrams for obtaining a frame sequence corresponding to a video frame provided in an embodiment of the present application;

[0038] Figure 6 This is the second example diagram of obtaining a frame sequence corresponding to a video frame provided in an embodiment of the present application;

[0039] Figure 7 Figure 3 of the principle example of splitting video segment a provided in an embodiment of the present application;

[0040] Figure 8 This is an example diagram of a video segment a that can be filtered out provided in an embodiment of the present application;

[0041] Figure 9 This is one of the example images of a video frame including checking the contents of a packaging box provided in an embodiment of the present application;

[0042] Figure 10 This is a second example of a video frame including checking the contents of a packaging box provided in an embodiment of the present application;

[0043] Figure 11 This is a third example of a video frame including checking the contents of a packaging box provided in an embodiment of the present application;

[0044] Figure 12 This is an example diagram of the segmented video segment a provided in an embodiment of the present application;

[0045] Figure 13 This is an example diagram of the segmented video segment b provided in an embodiment of the present application;

[0046] Figure 14 A flowchart of a packing detection method provided in an embodiment of the present application;

[0047] Figure 15 An example diagram of a display of an electronic device provided in an embodiment of the present application;

[0048] Figure 16 This is an example diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0050] With the development of technology, electronic products (such as mobile phones) have become an important part of people's lives. The production and packaging processes of various electronic products are now automatically completed by production lines. For example, the equipment production line is used to produce electronic products, and the packaging production line is used to package electronic products.

[0051] The packaging process involves placing the electronic product in a box and then sealing the box with plastic film or affixing a seal. The box, plastic film, and seal used to package the electronic product can all be referred to as the outer packaging of the electronic product.

[0052] Currently, packaging production lines are capable of automatically packaging electronic products using robotic equipment. However, even the most precise machines can make mistakes during operation. For example, boxes can be crushed, plastic sealing can be improperly applied, plastic film can be damaged, or seals can be misplaced. Furthermore, machines cannot detect defects in packaging materials like boxes, plastic film, and seals before packaging. Consequently, electronic products packaged on packaging production lines occasionally experience damaged outer packaging. It's easy to imagine that if electronic products with damaged outer packaging were shipped directly from the factory, consumers would question the product's quality control.

[0053] Therefore, all electronic products that pass through the packaging line are manually inspected by an operator. After confirming that the packaging is intact, the operator places the electronic product into a box (such as a cardboard box). Once the box is filled with the inspected electronic products, it is sealed.

[0054] However, during manual inspections of electronic product packaging, operators often perform inaccurate inspections. This can lead to damaged packaging being mistaken for intact products and thus directly packed. Furthermore, boxes may be packed before they are fully packed, or boxes may be forgotten to be sealed. These problems are all caused by unsatisfactory packaging inspections. Furthermore, if these issues are not promptly identified and corrected, subsequent rework will increase. For example, boxes containing defective or partially packed electronic products must be identified and repacked for re-inspection. Furthermore, if packaging defects or underfilled electronic products are not promptly identified before shipment, customer dissatisfaction may result. Therefore, these issues need to be managed. Manual management often wastes human and material resources and is prone to errors.

[0055] In order to improve the above problems, the present invention provides a packaging detection system. Figure 1 As shown, the above-mentioned box inspection system may include a camera 101 and a display device 102. The camera 101 may be mounted in a fixed position, with its field of view covering the boxing operation platform. The platform may contain boxes for electronic products whose outer packaging has been inspected by an operator. Furthermore, the camera 101's field of view may also include the operator, so that the camera 101 can capture images of the operator inspecting the outer packaging of the electronic products and placing the electronic products into the boxes.

[0056] In addition, the camera 101 is connected to the display device 102 for communication, and the camera 101 can send the video stream collected in real time to the display device 102. In this way, the display device 102 can also display the received video stream in real time.

[0057] In some embodiments, the display device 102 may be an electronic device with display and computing capabilities. Specifically, the display device 102 may not only display the video stream sent by the camera, but also analyze, based on the video stream, whether the operator's actions of inspecting the outer packaging of the electronic product (hereinafter referred to as the inspection action) meet preset standards.

[0058] In other embodiments, the display device 102 may also be an electronic device that only has a display function. In this way, after capturing the video stream, the camera 101 can analyze the operator's inspection actions based on the video stream to see if they meet the preset standards, and send the analysis results and the video stream to the display device 102, instructing the display device 102 to display the video stream and the corresponding analysis results. Alternatively, the camera 101 can send the video stream to the display device 102 and a third-party device (cloud server), and the third-party device can analyze the operator's inspection actions based on the video stream to see if they meet the preset standards. The third-party device then sends the obtained analysis results to the display device 102. In this way, the display device 102 can also display the video stream and the corresponding analysis results.

[0059] For example, Figure 1 As shown, the display device 102 may be a smart large screen. Of course, in other possible embodiments, the display device 102 may also be a mobile phone, a tablet computer, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an augmented reality (AR) or virtual reality (VR) device, etc. It is understandable that the embodiment of the present application does not impose any particular limitation on the specific form of the display device 102.

[0060] Please refer to Figure 2 , which is a structural diagram of a display device (also referred to as electronic device 100) provided in an embodiment of the present application.

[0061] like Figure 2 As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0062] Among them, the above-mentioned sensor module 180 may include sensors such as pressure sensor, gyroscope sensor, air pressure sensor, magnetic sensor, acceleration sensor, distance sensor, proximity light sensor, fingerprint sensor, temperature sensor, touch sensor, ambient light sensor and bone conduction sensor.

[0063] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0064] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0065] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0066] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0067] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0068] It is understood that the interface connection relationship between the modules illustrated in this embodiment is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0069] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0070] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLED, or a quantum dot light-emitting diode (QLED).

[0071] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0072] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0073] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include N cameras 193, where N is a positive integer greater than 1.

[0074] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0075] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0076] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0077] In the embodiment of the present application, a packing detection method is also provided, which can be applied to Figure 1 The box inspection system shown.

[0078] For example, the process of executing the above method by the above packaging inspection system may include: a camera capturing a video stream, then the camera and a display device collaboratively processing the video stream, and finally, the display device displaying the processed video stream.

[0079] The camera and display device collaborate to process the video stream, including:

[0080] After the camera captures the video stream, the video stream is preprocessed. For example, the camera can perform target detection on each frame of the captured video stream, and after detecting the target object (or object of interest), mark the detection window (or external frame) of the target object in the video frame. It can be understood that the above-mentioned target object can be all objects in the video frame except the background, or the above-mentioned target object can also be a pre-specified object (such as a box, a packaging box, an operator, etc.). In addition, the specific implementation process of the above-mentioned target detection can refer to the description in the subsequent embodiments. For another example, the camera identifies and marks video clip a from the video stream, and the video clip a records the action of the operator checking the outer packaging of the electronic product. Of course, the implementation details of the above-mentioned identification of video clip a can refer to the detailed description in the subsequent embodiments.

[0081] In addition, the above-mentioned camera and display device collaboratively process video streams including:

[0082] After the display device receives the video stream pre-processed by the camera, the display device analyzes whether the operator's detection action meets the preset standard based on the video segment a. The analysis result can be called the outer packaging inspection result of the inspected electronic product.

[0083] In other embodiments, the above method may also be applied to the display device in the above-mentioned box inspection system.

[0084] For example, after a display device receives a video stream captured by a camera, the display device can process the video stream and display the processed video stream. It is understood that the process of processing the video stream can include: performing target detection on the video frames in the video stream, and after detecting the target object, marking the corresponding detection window in the video frame. Then, each time a video segment a is identified, the display device analyzes the operator's inspection actions based on video segment a to see whether they meet preset standards. The display device then displays the corresponding analysis results. In this way, if the analysis results indicate that the operator's inspection actions for the electronic product packaging are not standardized, the display device can promptly remind the operator to re-inspect.

[0085] Of course, the details of the actual processing of the video stream in the above two embodiments are the same, the difference is that different processing steps are completed by different devices. Below, with reference to the accompanying drawings, the details of the method provided in the embodiment of the present application are described, taking the electronic device (e.g., display device 102) independently processing the video stream as an example.

[0086] In some embodiments, the operator checks the outer packaging of each electronic product in turn within the field of view of the camera, and places the electronic products whose outer packaging is confirmed to be intact into a box. In this way, the camera can record the entire process of the operator checking the outer packaging of the electronic products by collecting video streams, and send the collected video streams to the electronic device (such as a display device).

[0087] In some embodiments, after the electronic device receives the video stream, the electronic device may perform target detection on each video frame in sequence according to the acquisition time sequence of the video frames in the video stream.

[0088] In some embodiments, the electronic device may use a target detection model to identify target objects appearing in a video frame. The target detection model may be a neural network model trained using a target detection algorithm. During the training of the target detection model, boxes, packaging boxes, the operator's left hand, the operator's right hand, and the like in the training sample image may be marked as target objects. In this way, the target detection model has the ability to identify target objects such as boxes, packaging boxes, the operator's left hand, and the operator's right hand. That is, after the target detection model identifies a video frame containing a box, the output result obtained indicates that the target object includes a box. Additionally, the output result may also include a detection window corresponding to the target object "box," which is marked in the video frame. The detection window may be an external frame of the target object in the video frame, used to demarcate the image area occupied by the target object in the video frame.

[0089] As an example, the above-mentioned object detection model can be configured in an electronic device. In this way, after receiving a video stream, the electronic device can traverse the video frames in the video stream in the order in which they were acquired, and then use the object detection model to process the traversed video frames to obtain corresponding object detection results.

[0090] As an example, the above-mentioned target detection model can also be configured on a cloud server. In this way, after the electronic device receives the video stream, it can traverse the video frames in the video stream in the order of acquisition, and then send the traversed video frames to the cloud server. The cloud server uses the target detection model to process the video frames and obtain the target detection results corresponding to the video frames. The cloud server then feeds the target detection results of the video frames back to the electronic device.

[0091] The following mainly takes the scenario where the target detection model is configured in an electronic device as an example. The electronic device processes the video frame f through the target detection model. t After that, the target detection result can be obtained, that is, the detection window set Among them, the above O t Indicates video frame f t The corresponding detection window set. is the video frame f t The mth target object in the video frame f t The basic information in , for example, can be (x m ,y m ) indicates the center coordinates of the detection window (also called the bounding box) corresponding to the mth target object, (w m , h m ) indicates the length and width of the detection window corresponding to the mth target object, c m Indicates the number of target object types contained in the detection window corresponding to the mth target object. It can be understood that in the video frame f t In the video frame f, some target objects may overlap. For example, t In the box, there is a packaging box placed in the box, then the detection window corresponding to the box contains two types of target objects, the box and the packaging box. At this time, the c of the detection window corresponding to the box m The value is 2. In addition, m can take any integer value between 1 and M, where M is the video frame f t The total number of target objects in .

[0092] In some embodiments, the electronic device may further determine the video segment a from the video stream according to the target detection results corresponding to each video frame.

[0093] In some embodiments, the object detection results corresponding to the video frames may indicate the type and number of target objects within the camera's field of view. Thus, the electronic device can determine changes in the target objects within the camera's field of view, such as the presence or absence of new target objects, based on the target objects in multiple adjacent video frames.

[0094] In some embodiments, when a new packaging box (target object) appears in the camera's field of view, it means that the operator needs to start inspecting the packaging box. It is understandable that in the embodiments of this application, the packaging box mentioned can be a packaging box that already contains electronic products. In some scenarios, the packaging box and the outer packaging of the electronic product can be used interchangeably. For example, the above-mentioned inspection of the packaging box can also be referred to as the inspection of the outer packaging of the electronic product.

[0095] Furthermore, once the operator places the packaging box into the box, it signifies that the operator has completed the inspection of the electronic product's outer packaging. Thus, the video clip recording the moment the packaging box enters the camera's field of view and is placed into the box can be used to determine whether the operator's inspection of the electronic product's outer packaging is in compliance with regulations, i.e., it can be used as video clip a. It can be understood that the operator completes the inspection of the electronic product's outer packaging within the camera's field of view. Thus, if the camera does not capture the operator's inspection of the electronic product's outer packaging from the moment the packaging box enters the camera's field of view to the moment it is placed into the box, then it is deemed that the operator has not inspected the packaging box. If the camera captures the operator's inspection of the electronic product's outer packaging, then based on the corresponding video clip a, it can be determined whether the operator's inspection met the preset standards.

[0096] As an implementation method, the electronic device can determine video segment a from the video stream based on the target detection results corresponding to each video frame (e.g., whether a package box to be inspected appears and the location of the package box to be inspected in the video frame). The package box to be inspected can be a package box in the operator's hand and not placed in the box.

[0097] The first implementation method, such as Figure 3 As shown, based on the target detection result corresponding to video frame A, the electronic device can determine that video frame A includes box 301, packaging box 302, packaging box 303 and operator 304. Among them, packaging box 302 and packaging box 303 both contain electronic products (such as mobile phones). In addition, packaging box 302 and packaging box 303 have been placed in box 301, which means that the operator has already checked the packaging boxes. Figure 3 As shown, based on the object detection results corresponding to video frame B, the electronic device can determine that video frame B includes box 301, packaging box 302, packaging box 303, packaging box 305, and operator 304. Similar to video frame A, packaging box 302 and packaging box 303 are both located within box 301. Unlike video frame A, video frame B has a new packaging box 305, which also contains an electronic product. Video frame B shows that packaging box 305 follows the operator's hand into the camera's field of view, indicating that it is the packaging box to be inspected by the operator and can be referred to as the packaging box to be inspected. If video frame A and video frame B are adjacent video frames, video frame B can be determined to be the starting frame of video segment a and marked as such.

[0098] Afterwards, if Figure 4As shown, the electronic device can track the displacement of packaging box 305 based on the video frame whose acquisition time is after video frame B. When it is detected that packaging box 305 has been placed in box 301, that is, video frame C is detected. According to the target detection result of video frame C, the electronic device can determine that the center position of packaging box 305 in video frame C is within the detection window of box 301, that is, video frame C can indicate that packaging box 305 has been placed in box 301. In this way, the electronic device can determine video frame C as the end frame of video segment a and mark it.

[0099] After each set of start and end frames is determined, Figure 4 As shown, the electronic device can determine all video frames between video frame B and video frame C as the video frames corresponding to video segment a. Of course, the operator can then put other packages into box 301 in sequence. In this scenario, the electronic device can determine multiple video segments a from the video stream, each of which corresponds to a package to be inspected. For example, Figure 4 The video segment a shown corresponds to the packaging box to be inspected (packing box 305 ). The electronic device can determine whether the operator meets the preset standards when inspecting the packaging box 305 based on the video segment a.

[0100] Of course, other scenarios may arise during the operator's actual inspection of electronic product packaging. For example, if a defect is detected in the packaging box, the operator may move the defective packaging box out of the camera's field of view, or may not place the defective packaging box in the box. In this way, based on the video stream, the electronic device can track that packaging box 305 was not placed in box 301, but was instead moved out of the camera's field of view. In this scenario, the electronic device can clear the start frame marker corresponding to video frame B and abandon the acquisition of video segment a. The acquisition of video segment a will resume the next time the packaging box to be inspected is detected in the video stream.

[0101] In a second implementation, after obtaining the target detection result of each video frame in the video stream, each video frame is traversed in the order of acquisition time.

[0102] like Figure 5 As shown, when traversing to video frame A, the corresponding frame sequence A is obtained. Frame sequence A includes video frame A and also includes video frames whose acquisition time is adjacent to video frame A, such as video frame B. Exemplarily, in frame sequence A, the number of video frames whose acquisition time is before video frame A is the same as the number of video frames whose acquisition time is after video frame A. Of course, in other examples, the number may be different.

[0103] As an implementation method, the above frame sequence can be recorded as Indicates that the frame sequence includes the nth frame in the video stream. n can be any integer between t and t+h. t is greater than h, and h is an empirical value and can be a positive integer. t is the sequence number of the traversed video frame, also a positive integer. For example, if h is 2 and video frame A is the tth frame in the video stream, frame sequence A includes frames t, t-1, t-2, t+1, and t+2 in the video stream.

[0104] The electronic device then determines the identifier corresponding to video frame A based on the object detection results corresponding to each video frame in frame sequence A. This identifier can be either 0 or 1. If the number of frames in frame sequence A that do not contain the package to be inspected exceeds a threshold of 1, video frame A is labeled 0; otherwise, it is labeled 1. This threshold of 1 is an empirical value.

[0105] like Figure 5 As shown, it is assumed that the value of h is 2 and the value of threshold 1 is 3. In the frame sequence A, neither the video frame A nor the video frame whose acquisition time is before the video frame A contains the packaging box to be checked, while the video frame B and the video frame whose acquisition time is after the video frame B both contain the packaging box to be checked. Then the number of frames in the frame sequence A that do not contain the packaging box to be checked is 3, and the number of frames that do not contain the packaging box to be checked exceeds the corresponding threshold 1. It can be determined that the corresponding identifier of video frame A is 0. It should be noted that the above h value of 2 and the threshold 1 value of 3 are only used to illustrate the principles of "obtaining a frame sequence" and "determining the identifier of each video frame based on the frame sequence", and do not represent the reasonable values ​​used in actual applications. As shown in the above embodiments, the above h and threshold 1 are both empirical values, which can be the average values ​​obtained by statistics based on the intervals between the operator's inspection of the packaging box. That is, for different operators, the enabled h and threshold 1 can be different.

[0106] In addition, the above-mentioned packaging boxes to be inspected refer to packaging boxes that are not placed in boxes and are held in the hands of operators, such as Figure 3 The package box 305 in the figure may be a package box to be inspected, while the package boxes 302 and 303 are not package boxes to be inspected.

[0107] In some embodiments, video frame B is the next adjacent frame of video frame A. After determining the identifier corresponding to video frame A, the electronic device can traverse to video frame B. Figure 6 As shown, the frame sequence corresponding to video frame B is obtained from the video stream, that is, frame sequence B. If the value of h is 2 and the value of threshold 1 is 3. In frame sequence B, video frame B and the video frames acquired after video frame B both contain the packaging box to be inspected, while video frame A and the video frames acquired before video frame A do not contain the packaging box to be inspected. Therefore, the number of frames in frame sequence B that do not contain the packaging box to be inspected is 2, which is less than the corresponding threshold 1. Therefore, it can be determined that the identifier corresponding to video frame B is 1.

[0108] In this way, after determining the identifier of each video frame, if the identifier of the previous video frame is 0 and the identifier of the current video frame is 1, then the current frame is determined to be the starting frame of video segment a; if the identifier of the previous video frame is 1 and the identifier of the next video frame is 0, then the current frame is determined to be the ending frame of video segment a. For example, Figure 7 As shown, the previous frame of video frame B corresponds to an identifier of 0, and the corresponding identifier of video frame B is 1, then video frame B is determined to be the starting frame of video segment a. The previous frame of video frame C corresponds to an identifier of 1, and the corresponding identifier of video frame C is 0, then video frame C is determined to be the ending frame of video segment a. In the case that the starting frame closest to video frame C in acquisition time is video frame B, it can be determined that video segment a includes the video frames between video frame B and video frame C. In addition, it can be understood that Figure 7 The number of frames of the video segment a shown is only an example. The number of frames of the video segment a obtained in actual situations may be more or less, and this embodiment of the present application does not limit this.

[0109] Each video segment a obtained by the above implementation method also corresponds to a packaging box to be inspected. In this way, after the electronic device determines a video segment a, the electronic device can track the position change of the corresponding packaging box to be inspected in the video segment a. If it is determined through tracking that the packaging box to be inspected is finally placed in the box, for example, the inspection window of the packaging box to be inspected is tracked to enter the detection window of the box, then it is determined that the video segment a is usable for subsequent analysis and it is saved. If it is determined through tracking that the packaging box to be inspected moves out of the field of view of the camera, then it is determined that the video segment a is not suitable for subsequent analysis and can be discarded. For example, Figure 8 As shown, the target detection result of video frame D indicates that the packaging box to be inspected (i.e., packaging box 801) enters the field of view of the camera, and the target detection result of video frame E indicates that the packaging box to be inspected (i.e., packaging box 801) moves out of the field of view of the camera, then it is determined that the video segment a composed of the video frames between video frame D and video frame E is not suitable for subsequent analysis and can be screened out.

[0110] In some embodiments, the electronic device can also identify whether the operator's actions comply with the specifications based on the video segment a.

[0111] For example, the electronic device may be pre-configured with a motion recognition model. This motion recognition model may be a machine learning model trained using any algorithm, such as SlowFast, X3D, or TSM, combined with multiple sample videos. The training process of this motion model can be referenced in related technologies and will not be elaborated here.

[0112] In addition, the above-mentioned sample videos include a first type of video and a second type of video. The first type of video records the process in which the operator completes the outer packaging inspection according to the preset standard, that is, the video of the qualified inspection action. The second type of video records the process in which the operator fails to complete the outer packaging inspection according to the preset standard, that is, the video of the unqualified inspection action. In this way, the action recognition model can identify whether the inspection action of the operator in the video segment a is qualified based on the video segment a. For example, the video segment a is input into the action recognition model. If the action recognition model outputs a first value (such as 1), it indicates that the action of the operator checking the outer packaging in the video segment a meets the preset standard, that is, the outer packaging inspection result for the packaging box to be inspected is qualified. If the action recognition model outputs a second value (such as 0), it indicates that the action of the operator checking the outer packaging in the video segment a does not meet the preset standard, that is, the outer packaging inspection result for the packaging box to be inspected is unqualified.

[0113] In other embodiments, the electronic device can also determine whether the operator rotates the package box to be inspected and whether each side of the package box to be inspected faces the operator and stays for a specified time based on the target detection results in the video frames in the video segment a. Figure 9 、 Figure 10 and Figure 11 In the video frame shown, the operator rotates the package box to be inspected (package box 305) so that the six sides of the package box 305 are rotated in sequence to face the operator's eyes. If this is the case, it can be determined that the operator's inspection process for the package box 305 meets the set standards. Otherwise, it is determined that the set standards are not met.

[0114] If the outer packaging inspection result corresponding to video clip a fails to meet the set standards, the electronic device can display a corresponding warning message to remind the operator to re-inspect the packaging box in video clip a. In addition to displaying the warning message, the electronic device can also promptly remind the operator to re-inspect the packaging box by broadcasting a voice prompt or sending a warning message to other designated devices.

[0115] In summary, in this embodiment of the present application, the electronic device uses an object detection model to identify a target object in a video frame and tracks its position. Based on the tracking results, video segment a can be identified. After obtaining video segment a, the electronic device then uses an action recognition model to determine whether the operator's inspection of the outer packaging in video segment a is acceptable.

[0116] In the related art, the electronic device only inputs the entire video stream into a set model, and the set model identifies whether the operator's inspection process of each packaging box is qualified.

[0117] Compared to related technologies, both object detection and action recognition models are simpler and more lightweight than pre-defined models. However, in actual operation, cameras capture a variety of complex inspection scenarios (for example, placing a package into a box and then taking it out). If the sample data used to train the pre-defined model doesn't include such complex inspection scenarios, the pre-defined model will be unable to make judgments based on them. In such scenarios, the pre-defined model will need to be retrained, which undoubtedly increases the cost of maintaining the pre-defined model.

[0118] However, this application only needs to configure simple judgment logic in addition to the models used (target detection model and action recognition) to cope with various complex inspection scenarios. For example, the electronic device only needs to analyze whether the inspection process of the operator in the video clip a is compliant to obtain the corresponding outer packaging inspection results. If the operator takes out any package in the box, that is, after recognizing that the inspected package is removed from the box, the outer packaging inspection result for the package is canceled.

[0119] In some embodiments, the electronic device can also determine the number of inspected packaging boxes in the box based on the target detection results of the video frame. If the number of packaging boxes in the box reaches a predetermined value, then it is determined whether the target detection result corresponding to the video frame contains the drawer strap located in the box. If the box does not contain a drawer strap, the electronic device displays a corresponding warning message. The corresponding warning message can show the problem that the operator did not put the drawer strap into the box. For example, it can be the text "Drawer strap not detected, please replay", which can remind the operator to put the drawer strap into the box. In addition, in addition to displaying the corresponding warning message, the electronic device can also promptly remind the operator to replay the drawer strap by broadcasting a voice prompt or sending a warning message to other designated devices.

[0120] For example, Figure 12 As shown, the predetermined value corresponding to box 301 is 8. If, in video segment a, which is comprised of the video frames between video frame F and video frame G, there is at least one video frame showing that the number of packaging boxes in box 301 is equal to 8, then the video frames in video segment a are examined to determine whether the drawer strap has been placed in box 301. For example, if it is determined that video frame G contains drawer strap 1201 and that drawer strap 1201 is located between any two adjacent packaging boxes in box 301, the determination described in the subsequent embodiments is continued. Otherwise, a corresponding warning message is displayed to remind the operator to place the drawer strap in the box.

[0121] In some embodiments, the electronic device also needs to determine whether the box's C-surface appears in the video frame. If the box is placed on an operating table, the C-surface may be the outer surface opposite the box's bottom surface. The bottom surface of the box is the outer surface that directly contacts the operating table. Furthermore, the C-surface is an outer surface that can only be detected after the box is sealed. That is, detecting the C-surface indicates that the operator has completed the sealing process.

[0122] For example, when the electronic device detects a C-plane in a video frame, if the number of packages in the box has not reached a set value, the electronic device can display a corresponding warning message to inform the operator that the box is not yet full. This corresponding warning message can indicate that the operator has not fully filled the box, for example, it can be a text message such as "Please note, the box is not yet full." In addition to displaying the corresponding warning message, the electronic device can also promptly remind the operator by broadcasting a voice prompt or sending a warning message to a designated other device.

[0123] If the number of boxes in a box reaches a set value but the box does not contain a drawstring, the electronic device can display a corresponding warning message, prompting the operator to put the drawstring into the box and then repack the box.

[0124] In some embodiments, as Figure 13 As shown, the electronic device detects the C side of the box in the video frame H of the video stream, that is, the outer surface 1301. When the electronic device determines that the target detection result of the video frame H includes the C side, if it has been determined that the number of packaging boxes in the box is equal to the set value, and the box contains a drawer tape, then the electronic device can obtain the video segment b. The starting frame of the above-mentioned video segment b is the video frame H. After the video frame H is detected, the position change of the box is tracked, and when the B side of the box is tracked to appear, that is, when it is determined that the video frame I contains the B side of the box, the video frame I is used as the end frame corresponding to the video segment b. In this way, it is determined that the video frames between the video frame H and the video frame I constitute the video segment b. The B side of the box is an outer surface different from the C side, and a QR code is affixed to the B side. The QR code can be a QR code used for scanning and registration after the packaging is completed. As shown Figure 13 As shown, surface B may be an outer surface 1304 on which a QR code 1303 is affixed.

[0125] The electronic device then inputs video clip b into the action recognition model, which determines whether the operator's packaging process is satisfactory. For example, the action recognition model determines whether the operator has sealed the box as required and whether the seal has been affixed. Seals are also identifiable targets for the object detection model, such as seal 1302 in video frame H.

[0126] As a way to implement Figure 14As shown, the electronic device cuts out a video segment a containing an operator inspecting the outer packaging of an electronic product from the video stream according to the target object type and the position of each target object contained in the video frame in the video stream. The specific implementation details can be referred to the description in the aforementioned embodiment and will not be repeated here.

[0127] The electronic device uses a motion recognition model to analyze video clip a and assess whether the operator's inspection of the outer packaging meets preset standards. For example, consider a video clip a recording of an operator inspecting package a. The motion recognition model determines whether the operator's inspection of package a meets the required standards, for example, whether they have thoroughly inspected all exterior surfaces of package a.

[0128] If the motion recognition model determines that the operator's inspection process of package a does not meet the standards, the electronic device will display the outer packaging inspection result of package a as unqualified, prompting the operator to recheck. In this scenario, the operator is required to take package a out of the box and recheck it. That is, after the electronic device detects that package a follows the operator's hand out of the box, the electronic device cancels the display of the outer packaging inspection result of package a. Of course, after the operator completes the inspection of package a again, package a can be put into the box again. In this way, the electronic device can again cut out a video clip a for package a from the video stream and use the motion recognition model to make a judgment again. When the motion recognition model determines that the operator's inspection process of package a is qualified, the electronic device displays the outer packaging inspection result of package a as qualified.

[0129] If the motion recognition model determines that the operator's inspection of package a meets the standards, the electronic device can obtain the total number of packages in the box. Furthermore, the electronic device can display the inspection result of the outer packaging corresponding to package a as qualified, and remind the operator to continue inspecting the outer packaging of the electronic product.

[0130] Furthermore, if the operator removes box a from the box after the outer packaging inspection result has been determined to be acceptable, the outer packaging inspection result for box a will remain as long as box a remains within the camera's field of view. This means that after removing box a, the operator can return it to the box without re-inspecting it. If it does leave the camera's field of view, the outer packaging inspection result for box a is canceled, and the total number of boxes in the box is recounted.

[0131] After determining the total number of packaging boxes in the box, the electronic device can determine whether the box is full based on the total number of packaging boxes in the box, for example, determine whether the total number of packaging boxes reaches a set value corresponding to the box.

[0132] If the box is not full (the total number of boxes is less than the set value), the system continues to check the video stream for the next box (e.g., box b) and checks if video clip a appears. The system then determines if the operator's inspection of box b has passed in the same manner and displays the corresponding outer packaging inspection result. Additionally, if the total number of boxes in the box exceeds the set value, the electronic device can display a warning message, alerting the operator that there are too many electronic products in the box.

[0133] If the box is full (the total quantity is equal to the set value), a video frame captured later than video segment a is used to detect whether a drawer strap is present in the box. If no drawer strap is present in the box, a reminder message 1 is displayed to prompt the operator to place the drawer strap in the box.

[0134] If a drawstring is present in the box, the operator can continue to determine whether the box is packed according to the preset standards based on the video frames in the video stream. For example, video segment b can be segmented based on the video frames showing side C and side B. Then, the action recognition model can be used to determine whether the operator packed the box according to the preset standards.

[0135] If the motion recognition model determines that the operator's packaging process is unsatisfactory (for example, the box is not plastic-sealed or the seal is not affixed), the electronic device can display a reminder message 2 to prompt the operator to repack the box. If the motion recognition model determines that the operator's packaging process is satisfactory, the process ends.

[0136] In addition, if Figure 15 As shown, the electronic device can not only display the video stream captured by the camera in real time, but also display the outer packaging inspection results in a split screen. In the outer packaging inspection results, the object ID can be a serial number automatically assigned by the electronic device to the target object. The corresponding object ID can be assigned according to the order in which the packages enter the box. For example, if package 302 enters the box first, the corresponding object ID may be 1. Package 303 enters the box after package 302, the corresponding object ID may be 2. Package 305 enters the box after package 303, and the corresponding object ID may be 3.

[0137] In addition, the outer packaging inspection results displayed by the electronic device can be updated in real time. For example, if the electronic device determines that the operator's inspection process of packaging box 305 is unqualified, then the corresponding packaging box (object ID is 3) will be displayed with an unqualified mark. In this case, if the operator removes packaging box 305 and re-inspects packaging box 305 according to the preset standards, the electronic device can detect that the operator's inspection process of packaging box (object ID is 3) is qualified, and the electronic device can change the unqualified mark for object ID 3 to a qualified mark.

[0138] The following takes the outer packaging inspection of the first product as an example to describe the scenarios that may be encountered in actual applications of the solution provided by the embodiment of the present application.

[0139] The first product may be an electronic product or other types of products (e.g., clothing, shoes, hats, etc.). The electronic device may receive a video stream captured by a camera in real time, which may be divided into a first video stream, a second video stream, a third video stream, etc., according to different time periods corresponding to the video streams.

[0140] When the first video stream is received, if a video segment a (also referred to as the first segment) for the first product is identified from the first video stream, the electronic device evaluates whether the operator's first action meets the first standard (that is, the aforementioned preset standard) based on the first segment and the preconfigured first model (that is, the aforementioned action recognition model). The above-mentioned first standard can be pre-set by the user. For example, the first standard can be that the operator needs to rotate the product during the inspection of the product's outer packaging, inspect each outer surface of the product in turn, and the dwell time on each outer surface reaches a specified time length. Of course, the first standard can also be other action steps for inspecting the outer packaging set by the user.

[0141] The first clip above records the process from the operator bringing the first product into the first space (the spatial area within the camera's field of view) to placing the first product into the first box (a box within the camera's field of view, which can be used to store inspected electronic products). The first action recorded in the first clip is the action of inspecting the outer packaging of the first product.

[0142] Specifically, the first segment can be identified by detecting that the first video frame is included in the first video stream, then tracking the position of the first product in the first video stream starting from the first video frame until a second video frame is detected. The first video frame is the first frame in which the first product appears, and the second video frame is the frame in which the first product is located within the first box. In this way, the video frames between the first and second video frames can be determined to constitute the first segment.

[0143] After determining the first segment, the first segment is input into the first model to obtain a first output result. Based on the value of the first output result, it is determined whether the first action meets the first standard. For example, if the value of the first output result is a first value, it indicates that the first standard is met, while if the value of the first output result is a second value, it indicates that the first standard is not met.

[0144] If the first model determines that the first action does not meet the first standard, a first prompt message may be displayed relative to the first identifier (e.g., object ID) of the first product, such as the text "Inspection Failed", prompting the operator to re-inspect the outer packaging of the first product. If the first model determines that the first action meets the first standard, a second prompt message may be displayed, such as the text "Inspection Passed", prompting the operator that the inspection of the first product is complete.

[0145] It can be understood that after the first product is placed in the first box, the electronic device analyzes whether the operator's inspection action on the first product meets the first standard. If it does not meet the requirements, the operator is instructed to re-check the first product by displaying the first prompt message. In this scenario, the operator will take the first product out of the first box and re-inspect it as prompted. The process of re-inspection will still be captured by the camera and transmitted to the electronic device. In this way, after the electronic device receives the second video stream (that is, the video stream collected after the first prompt message is displayed), it can identify whether the second video stream contains the second clip. The second clip also belongs to video clip a. Unlike the first clip, the second clip records the process of the operator taking the first product out of the first box and putting it back into the first box again. The second action is the action of re-inspecting the outer packaging of the first product recorded in the second clip.

[0146] The second video segment may be identified by tracking the position of the first product in the second video stream starting from the third video frame when the second video stream includes the third video frame until a fourth video frame is detected, wherein the center position of the first product in the third video frame does not overlap with the first box, the fourth video frame is a video frame in which the center position of the first product overlaps with the first box, and the third video frame is acquired earlier than the fourth video frame.

[0147] When the second action meets the first criterion, a second prompt message is displayed to prompt the operator that the inspection of the first product has been completed.

[0148] When the second prompt message is displayed, the first product has been placed in the first box. After the first product is placed in the first box, the electronic device can count the number of electronic products already in the first box, that is, the first number. Exemplarily, the electronic device can count the first number by obtaining the number of second detection windows contained in the first detection window based on a specified video frame. The first detection window can be an external frame corresponding to the first box, and the second detection window is an external frame corresponding to the electronic product, and the detection window of the first product is contained in the second detection window. In addition, the specified video frame can be a video frame in the first video stream whose acquisition time is after the first segment, or a video frame in the second video stream whose acquisition time is after the second segment. Specifically, if the first action in the first segment meets the first standard, the electronic device can determine the first number of electronic products in the first box after the first product is placed in the box based on the video frames in the first video stream. If the second action in the first segment meets the first standard, the electronic device can determine the first number of electronic products in the first box after the first product is placed in the box based on the video frames in the second video stream.

[0149] Of course, there will be two situations: the first quantity is equal to the first value (that is, the set value corresponding to the first box) and the first quantity is not equal to the first value:

[0150] In the first scenario, after determining that the first quantity is equal to the preset first value, it is also necessary to detect whether the first box contains the first drawer strap. For example, based on the video frame that determines the first quantity, it is identified whether the first box contains the first drawer strap. In the case of the first drawer strap, a third video stream captured by the camera is received (the video stream received after the second prompt message is displayed). If the third video stream contains a third segment, then based on the third segment and the first model, it is evaluated whether the operator's third action meets the second standard. The third segment records the operator's process of packing the first box, that is, the third segment contains a video frame with the C side of the first box, and also contains a video frame with the B side of the first box. The third action is the action of packing the first box recorded in the third segment. The above-mentioned second standard is also a preset standard, which is a standard for the action of packing the box and can be customized by the user. For example, in the process of packing the box, it is necessary to apply tape, seal, rotate and scan, etc.

[0151] In the case that the third action does not meet the second standard, a third prompt message is displayed (eg, text "packing action is unqualified, please repack") to prompt the operator to repack the first box.

[0152] In the second scenario, when the first quantity is less than the preset first value and a third video stream captured by the camera is received, if the third video stream also contains a third segment, a fourth prompt message is displayed, such as the text "The box is not packed", indicating that the first box is not full.

[0153] An embodiment of the present application further provides an electronic device, which may include a memory and one or more processors. The memory and processor are coupled. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device performs each step of the above embodiment. Of course, the electronic device includes but is not limited to the above memory and one or more processors.

[0154] The present application also provides a chip system, which can be applied to the terminal device in the above embodiment. Figure 16 As shown, the chip system includes at least one processor 2201 and at least one interface circuit 2202. The processor 2201 can be the processor in the above-mentioned electronic device. The processor 2201 and the interface circuit 2202 can be interconnected via a line. The processor 2201 can receive and execute computer instructions from the memory of the above-mentioned electronic device through the interface circuit 2202. When the computer instructions are executed by the processor 2201, the electronic device can perform the various steps in the above-mentioned embodiment. Of course, the chip system can also include other discrete components, which are not specifically limited in this embodiment of the present application.

[0155] In some embodiments, through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0156] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0157] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0158] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A packing detection method, characterized in that: Applied to an electronic device, the electronic device is communicatively connected to a camera, and the method includes: Receiving a first video stream captured by the camera; In response to detecting a first video frame in the first video stream, tracking the position of the first product in the first video stream starting from the first video frame until a second video frame is detected, wherein the first video frame is a video frame in which the first product first appears, and the first product in the second video frame is a video frame in which the first product is located in a first box; Determining that video frames between the first video frame and the second video frame constitute a first segment; Based on the first segment and a preconfigured first model, evaluating whether the operator's first action meets a first criterion; wherein the first model is a machine learning model for recognizing actions; the first segment records the process from the operator bringing a first product into a first space to placing the first product into a first box; the first space is the camera's acquisition space; the first box is used to store inspected electronic products; the first box is located in the first space; and the first action, recorded in the first segment, is the action of inspecting the outer packaging of the first product. If the first action does not meet the first standard, displaying a first prompt message corresponding to the first identifier of the first product, prompting the operator to recheck the outer packaging of the first product; receiving a second video stream captured by the camera; In response to the second video stream including a third video frame, tracking the position of the first product in the second video stream starting from the third video frame until a fourth video frame is detected, thereby obtaining a second segment, wherein the center position of the first product in the third video frame does not overlap with the first box, the fourth video frame is a video frame in which the center position of the first product overlaps with the first box, and the third video frame is captured earlier than the fourth video frame; evaluating, based on the second segment and the first model, whether a second action of the operator meets the first criterion, wherein the second segment records the process of the operator taking the first product out of the first box and putting it back into the first box, and the second action is the action of inspecting the outer packaging of the first product recorded in the second segment; If the second action meets the first criterion, display a second prompt message to prompt the operator that the inspection of the first product is completed; determining a first quantity based on video frames in the first video stream or the second video stream, where the first quantity is the number of the electronic products placed in the first box after the first product is placed in the first box; When the first quantity is less than a preset first value and a third video stream captured by the camera is received, if the third video stream includes a third segment, displaying a fourth prompt message indicating that the first box is not full, wherein the third segment records a process of the operator packing the first box; When the first quantity is equal to a preset first value and a third video stream captured by the camera is received, if the third video stream includes a third segment, then evaluating whether a third action of the operator meets a second criterion based on the third segment and the first model, wherein the third segment records a process of the operator packing the first box, and the third action is the action of packing the first box recorded in the third segment; In the case that the third action does not meet the second standard, a third prompt message is displayed to prompt the operator to repack the first box.

2. The method according to claim 1, characterized in that The method further comprises: When the first action meets the first standard, a second prompt message is displayed to prompt the operator that the inspection of the first product is completed.

3. The method according to claim 1, characterized in that After determining the first quantity, the method further includes: When the first quantity is equal to the first value, it is determined that the first box contains a first drawer strap based on the video frames in the first video stream or the second video stream.

4. The method according to claim 1, wherein The step of evaluating whether the operator's first action meets a first standard based on the first segment and the preconfigured first model includes: Inputting the first segment into the first model to obtain a first output result; Determine whether the first action meets the first criterion based on the value of the first output result.

5. An electronic device, characterized in that: The electronic device includes one or more processors and a memory; the memory is coupled to the processor, and the memory is used to store computer program code, the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the one or more processors are used to execute the method according to any one of claims 1 to 4.

6. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 4.

7. A computer program product, characterized in that The computer program product comprises a computer program which, when run on a computer, causes the computer to perform the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Packing method, packing scheme generation method, packing system and server

    CN108876230A

  • Packaging video recording method, device and system

    CN109672921A