Method and system for intelligently identifying and positioning small articles in home scene
By using the improved YOLOv8 target detection model and item position correlation description method in the home environment, combined with database management and voice broadcast, the precise identification and positioning of small items is achieved, the problem of item management in the home environment is solved, and the quality of life and efficiency are improved.
Patent Information
- Application Number
- CN202510236580.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-10
AI Technical Summary
In the home environment, item management is chaotic and it is difficult to quickly find and locate, resulting in wasting time and energy. Especially for special groups, it reduces the quality and happiness of home life.
A method and system for intelligent identification and positioning of small items in home scenes is adopted, indoor scene images are collected through cameras, and the improved YOLOv8 target detection model is used to combine the item position correlation description method to generate detailed item position information, and real-time positioning is achieved through database management and voice broadcasting.
It realizes accurate detection and positioning of small items, generates clear and intuitive item position information, solves the problems of insufficient detection accuracy and fuzzy location description in the existing technology, and improves object search efficiency and quality of life at home.
Smart Images

Figure CN120125807A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target object detection, and particularly relates to a method and system for intelligent recognition and positioning of small objects in a home scene. Background Art
[0002] In the busy and noisy modern life, the home environment carries our yearning for a quiet, comfortable and convenient life. The rationality of its design, the comfort of its layout, and the convenience of item management directly affect people's quality of life. However, with the accelerating pace of life and the increasing number of items, there are often problems of chaotic item management and difficulty in quickly finding and positioning items in home life. These problems not only bring a lot of inconvenience to the lives of ordinary people, such as frequently searching for lost items, wasting time and energy, but also pose huge challenges to the lives of special groups, and invisibly reduce the quality and happiness of home life. Therefore, solving the problems of item management and searching in the home environment is of great significance for improving the quality of life and creating a more convenient and comfortable home environment.
[0003] According to a research report "Orden y tiempo" conducted by Sigma Dos and published by IKEA, people spend almost about 5,000 hours in their lifetime looking for items at home. Nearly 60% of people say they spend 10 minutes a day looking for lost things, 20% say they spend nearly 20 minutes looking for items, and for people who often forget things, this time will be even more. For the general population, forgetting the specific location of items is a common phenomenon, but these seemingly trivial matters often bring inconvenience to people's lives. Especially when going out in a hurry in the morning, forgetting the location of the keys may lead to being late or missing an important meeting; when wanting to rest at night, not being able to find the air conditioner remote control may affect sleep and relaxation.
[0004] In view of the above situation, there is an urgent need for a method and system that can solve the problem of item recognition and positioning in the home environment, so as to improve people's home experience, enhance the convenience and comfort of life, and reduce the time and energy wasted due to frequently searching for items.
[0005] The detection and positioning technology of items is the key to solving the above-mentioned problem of item searching. However, most current methods cannot meet the requirements of precise search and positioning of items. For example, in 2019, Zhao Xiaojun et al. ("Design and Research of a Guide Stick Based on RFID / GPS" [J]. Video Engineering, 2019, 43(01): 111-114.) integrated GPS positioning technology and radio frequency identification technology (RFID) to develop an integrated auxiliary device. They used RFID tags for positioning and then used GPS for positioning and navigation. However, GPS is easily affected by the environment and has a large accuracy fluctuation. At the same time, the cost of laying and popularizing RFID tags is relatively high. In 2021, Mukhriddin M et al. ("Smart Glass System Using Deep Learning for the Blind and Visually Impaired" [J]. Electronics, 2021, 10(22): 2756-2756.) invented a pair of smart glasses that integrated functions such as GPS and cameras, which can perform positioning, object recognition, and audio feedback. However, the positioning is not clear, the cost is too high, and it is not convenient for popularization. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for intelligent identification and positioning of small items in a home scene, which solves the problem of inaccurate existing positioning methods.
[0007] The present invention is realized through the following technical solutions:
[0008] The present invention discloses a method for intelligent identification and positioning of small items in a home scene, including the following steps:
[0009] Step 1: Based on the division of the indoor detection area, collect frame images in the indoor scene for a period of time and preprocess the frame images;
[0010] The division of the indoor detection area is carried out during the first use. Specifically, the key detection areas of the indoor environment are divided and the names of the corresponding areas are marked to obtain the area information divided by the user. The area information divided by the user includes the coordinates of the area and its corresponding name;
[0011] Step 2: Send the preprocessed images into the improved YOLOv8 target detection model, detect each item and generate its detailed position information through the item position association description method, and perform position prediction on the occluded items to obtain the item position information;
[0012] Enter the item position information into the database for management;
[0013] Step 3: Listen to user instructions in real time. After identifying the user instructions, query the location information of the corresponding item in the database and announce it to the user through voice broadcast.
[0014] Furthermore, divide the key detection areas of the indoor environment and label the names of the corresponding areas to obtain the area information divided by the user. Specifically:
[0015] Capture the first frame image of the indoor environment as the background and display it on the operation interface.
[0016] The user manually frames and divides the detection area and names the area.
[0017] Automatically store the area information divided by the user.
[0018] Furthermore, in Step 1, capture frame images of the indoor scene for a period of time through the camera. Specifically:
[0019] The camera monitors the indoor scene in real time.
[0020] Use the frame difference method to compare adjacent frame images. When it is detected that the scene has changed significantly enough to affect the position of the item, start capturing frames of this period of time.
[0021] Furthermore, use the frame difference method to compare adjacent frame images. The specific processing process is as follows:
[0022] Convert the frame image to grayscale. The expression is:
[0023] Gray(x,y) = 0.299×R(x,y) + 0.587×G(x,y) + 0.114×B(x,y) (1)
[0024] Among them, R(x,y), G(x,y), and B(x,y) are the red, green, and blue color channel values of the color image at the pixel point (x,y), respectively, and Gray(x,y) is the grayscale value of the corresponding grayscale image at the pixel point (x,y); the frame image is a color image;
[0025] Calculate the difference between the grayscale images of adjacent frames. For each pixel point (x,y), the difference calculation expression is:
[0026] GrayDiff(x,y) = ∣Gray1(x,y) - Gray2(x,y)∣ (2)
[0027] Among them, Gray1(x,y) and Gray2(x,y) are the grayscale values of the two images at the pixel point (x,y), respectively, and GrayDiff(x,y) is the difference value of the corresponding difference image at the pixel point (x,y);
[0028] When the difference value is greater than the preset threshold, capture the current frame image.
[0029] Further, in step 1, preprocess the frame image, specifically:
[0030] According to the area information divided by the user obtained in step 1, cut the collected indoor scene frame images according to the coordinate positions of each area;
[0031] Name each small image obtained after cutting with the name corresponding to the area and store it.
[0032] Further, in step 2, send the preprocessed frame image into the improved YOLOv8 object detection model, detect each item and generate its detailed position information through the item position association description method, specifically:
[0033] Detect all the cut small images using the improved YOLOv8 object detection model;
[0034] Associate the items detected on each small image with the area corresponding to the small image and the reference objects existing in the area, and the position of the item is the relative position description of it with the reference object in the area;
[0035] The described item position association description method is:
[0036] Associate the preset area position and the fixed reference item with the target item to describe the actual position of the target item in reality.
[0037] Further, the improved YOLOv8 object detection model is: fuse the MSDA attention mechanism in the YOLOv8 model, capture richer context information by applying dilated convolutions at different scales, and dynamically adjust the importance of features through the MSDA attention mechanism;
[0038] Introduce the EfficientNetV1 backbone network to replace the original backbone structure of YOLOv8.
[0039] Further, in step 2, predict the position of the occluded item, specifically:
[0040] By comparing the detection results of the frame images within a period of time, when it is found that an item disappears from the scene image, it is initially considered that it is covered or taken out of the current scene;
[0041] Combined with the position coordinates of the disappeared item in the previous frame image, use the Kalman filter to predict its coordinate position after disappearance and judge its physical position in reality. The Kalman filter state equation is:
[0042] x_k = A × x_(k-1) + B × u_(k-1) + w_(k-1) (3)
[0043] Where x_k is the state vector at time k, A is the state transition matrix, x_(k-1) is the state vector at time k-1, B is the control input matrix, u_(k-1) is the control input at time k-1, and w_(k-1) is the process noise.
[0044] Furthermore, in step 2, the item location information is entered into the database for management, specifically:
[0045] The item location information includes the item name, the location area where it is located, the detection time, and whether it is covered.
[0046] The item location information is entered into the MySQL database, and steps 2 and 3 are repeatedly executed to achieve real-time update of the item location information in the database.
[0047] Step 3 is specifically:
[0048] Circularly monitor the user password information. When the user password is monitored, it is recognized as the command text.
[0049] The command text is transmitted to the database query program through serial communication. The database query program queries the location information of the target item in the database according to the command requirements.
[0050] The query result is used to broadcast the location information to the user in the form of voice.
[0051] The present invention also discloses a small item intelligent recognition and positioning system in a home scene, including:
[0052] An area division module, which is used for the first use to divide the key detection areas of the indoor environment and label the corresponding area names to obtain the area information divided by the user; the area information divided by the user includes the coordinates of the area and its corresponding name.
[0053] An image preprocessing module, which is used to preprocess the frame images in the indoor scene collected for a period of time.
[0054] A target detection module, which is used to send the preprocessed image into the improved YOLOv8 target detection model, detect each item, generate its detailed location information through the item location association description method, and perform location prediction on the occluded items to obtain the item location information.
[0055] A location information database, which is used to enter the item location information into the database for management.
[0056] A voice recognition module for recognizing real-time monitored user instructions;
[0057] A database query module for querying the location information of items stored in the location information database after recognizing user instructions and querying the location information of the corresponding items in the database;
[0058] A voice broadcast module for broadcasting the location information of the corresponding items queried from the database query module to the user by voice.
[0059] Compared with the prior art, the present invention has the following beneficial technical effects:
[0060] The present invention discloses a method and system for intelligent recognition and positioning of small items in a home scene. By adopting an optimized YOLOv8 object detection model and combining an innovative item location association description method, accurate detection of small items is achieved, and clear, intuitive and easy-to-understand item location information can be generated. Effectively solves the problems of insufficient detection accuracy and fuzzy location description existing in the current item detection and positioning technology.
[0061] In the improved YOLOv8 object detection model of the present invention, the EfficientNetV1 backbone network and the MSDA attention mechanism are first introduced into the YOLOv8 framework, and the effectiveness of this combination in the object detection task is verified through experiments. Although EfficientNetV1 and MSDA are prior arts, their combination in the application of YOLOv8 is new, and significantly improves the detection accuracy and context information capture ability of the original YOLOv8 model.
[0062] The MSDA attention mechanism captures multi-scale context information through dilated convolution and is more suitable for processing multi-scale targets; the multi-level attention mechanism (object level, pixel level, part level) of the comparative document is detailed, but not optimized for multi-scale targets and has a high computational complexity.
[0063] The basic model YOLOv8 is a one-stage model and has obvious advantages in terms of simplicity and speed compared to the traditional CNN (Convolutional Neural Network) based basic model in the document, and is more suitable for real-time detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0065] Figure 2 It is a schematic diagram of the database storage content of the present invention;
[0066] Figure 3 It is a schematic diagram of the network structure of the detection model of the present invention;
[0067] Figure 4 This is the working schematic diagram of the attention machine of the present invention;
[0068] Figure 5 This is the schematic diagram of the backbone network structure of the present invention;
[0069] Figure 6 This is the system working flow chart of the present invention. Detailed implementation manners
[0070] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0071] It should be noted that the terms "including" and "having" in the description and claims of the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0072] The present invention will be further described in detail below in conjunction with the accompanying drawings:
[0073] Embodiment 1
[0074] The present invention discloses a method for intelligent identification and positioning of small items in a home scene, including the following steps:
[0075] S1. When used for the first time, divide the key detection areas in the indoor environment and label the names of the corresponding areas;
[0076] Specifically, the key detection areas are the areas that users focus on or the areas where users often place items. For example, areas such as the coffee table area, the dining table area, and the sofa area can be manually framed by the user to select the detection areas.
[0077] Specifically, labeling the names of the corresponding areas is as follows: for example, after framing the coffee table area, name this area "coffee table area". This is done so that when the system locates an item and makes a voice broadcast later, it can broadcast the name of the area where the item is located.
[0078] S2. Image acquisition and processing: Collect frame images in the indoor scene for a period of time through a camera, and preprocess the collected frame images;
[0079] S3. Item Detection and Location: Send the preprocessed image into the improved YOLOv8 object detection model for detection, generate detailed location information of each item through the item location association description method, predict the location of occluded items, and enter all the obtained item location information into the database for management;
[0080] S4. Listen to user instructions in real time. When a search instruction is recognized, query the location information of the corresponding item in the database and announce it to the user through voice.
[0081] Embodiment 2
[0082] Details of S1 are introduced in detail. Before the system is used for the first time, it is necessary to divide and name the key detection areas of the indoor environment to be detected; specifically, it includes the following sub-steps:
[0083] S11. Call the camera to capture the first frame of the indoor environment image and display it as the background on the UI operation interface;
[0084] S12. Use the powerful Python GUI framework PyQt5 to build a graphical user interface (GUI). Integrate functions such as image display and button interaction in it, allow users to mark on the image. In addition, define the InputDialog dialog class to be used to input the area name after the user completes the area division;
[0085] S13. Use OpenCV to read the first frame image file of the indoor environment. When the user divides the detection area, capture the coordinates of the mouse click through the mouseReleaseEvent method of the ImageLabel class, and map these coordinates from the local image to the original image;
[0086] S14. Store the coordinates of each divided area and its corresponding name.
[0087] It should be noted that step S1 only needs to be performed when the system is used for the first time, and the division and naming results will be stored for subsequent use by the system.
[0088] Embodiment 3
[0089] Details of S2 are introduced in detail. S2 specifically includes the following sub-steps:
[0090] S21. The camera monitors the indoor scene in real time. The frame difference method is used to compare adjacent frame images. When it is detected that the scene has changed greatly enough to affect the item location, start to capture frames of the picture during this period. The reason is:
[0091] Since most of the items in the indoor environment are stationary most of the time, there is no need to repeatedly detect the items in this situation. Only when a change occurs in the current scene, and at this time the position of the items may change, then a new detection operation is required.
[0092] The frame difference method is used to compare adjacent frame images, and the specific processing process is as follows:
[0093] S211. Convert the video frame image to grayscale, and the expression is shown in formula (1):
[0094] Gray(x,y)=0.299×R(x,y)+0.587×G(x,y)+0.114×B(x,y) (1)
[0095] Where R(x,y), G(x,y), and B(x,y) are the red, green, and blue color channel values of the color image at the pixel point (x,y), respectively, and Gray(x,y) is the brightness value of the corresponding grayscale image at the pixel point (x,y).
[0096] S212. Calculate the difference between the grayscale images of adjacent frames. For each pixel point (x,y), the difference calculation expression is shown in formula (2):
[0097] GrayDiff(x,y)=∣Gray1(x,y)-Gray2(x,y)∣ (2)
[0098] Where Gray1(x,y) and Gray2(x,y) are the grayscale values of the two images at the pixel point (x,y), respectively, and GrayDiff(x,y) is the difference value of the corresponding difference image at the pixel point (x,y);
[0099] When the difference value is greater than the preset threshold, it is considered that a large change has occurred in the indoor environment and the position of the indoor items may change. A series of frame pictures during this period are intercepted at a certain time interval and stored;
[0100] When the scene picture stops changing after a period of time, stop intercepting the frame pictures and continue to monitor the indoor scene until the next change starts.
[0101] S22. When preprocessing the collected indoor scene frame images, the specific operation steps are as follows:
[0102] Crop the first and last frame pictures collected according to the area information divided by the user obtained in step S1, and crop the original frame image into small detection area images. This area is the detection area manually divided by the user, and the names of the cropped small images are named after their corresponding detection area names and stored.
[0103] The described image preprocessing method is innovative in that:
[0104] By operating in this way, each small image represents a known area of the user. When detecting objects in the small image, any object detected in the small image can be considered to exist in that area.
[0105] The traditional object detection and positioning results are represented in the form of coordinates, which are not easily understood by users. By operating in this way, the object can be associated with the known area, generating a description statement of the relative position of the object that is easy for users to understand, thus improving the efficiency of finding objects.
[0106] Moreover, when the original large image is cropped into small images and then object detection is performed, the detection accuracy of the small target objects by the target detection model can be improved, the detection rate of objects can be increased, and the false detection rate can be reduced.
[0107] Example 4
[0108] Details of S3 are introduced. Step S3 specifically includes:
[0109] S31. Detect the small images obtained by cutting the first frame image using the improved YOLOv8 target detection model;
[0110] S32. Use the object position association description method to associate the objects detected on each small image with the area corresponding to the small image and the reference objects existing in the area. The position of the object is the relative position description of it with respect to the area and the reference objects, and the detected objects and their positions are entered into the database.
[0111] The innovative point of the described object position association description method lies in:
[0112] Since the results of ordinary object positioning are represented in the form of coordinates, but this coordinate information is often not intuitive and easy for ordinary users to understand, lacking a detailed description of the environmental position of the object. Users usually need to spend extra time and effort to convert the coordinate information into the position in the actual environment, which reduces their efficiency of finding objects;
[0113] Therefore, the present invention adopts a method of associating the preset area position and fixed reference objects with the target object to describe the actual position of the target object in real life, facilitating user understanding and positioning and improving the efficiency of finding objects.
[0114] S33. According to the above steps, detect the small images obtained by cutting the last frame image using the improved YOLOv8 target detection model;
[0115] S34. Compare the detection results of the first frame image and the last frame image. When it is found that an item disappears from the scene picture, it is initially considered that it is covered or taken out of the current scene;
[0116] S35. For such disappearing items, use the method of target tracking and prediction to predict the coordinate position after their disappearance. The specific steps are as follows:
[0117] Utilize all the frame images collected previously to detect a series of coordinate trajectory information of the item before its disappearance, and combine with Kalman filtering to track and predict the approximate coordinate position after the item disappears;
[0118] The Kalman filter state equation is shown in formula (3):
[0119] x_k = A ×x_(k-1) + B ×u_(k-1) + w_(k-1) (3)
[0120] Where, x_k is the state vector at time k, A is the state transition matrix, x_(k-1) is the state vector at time k-1, B is the control input matrix, u_(k-1) is the control input at time k-1, and w_(k-1) is the process noise.
[0121] S36. Map the predicted coordinates to the real environment to obtain the position description information of the item in the real environment;
[0122] S37. Enter the item position information detected and predicted in the above steps into the database for management. Specifically:
[0123] Enter the detected item position information including detailed information such as the item name, location area, detection time, and whether it is covered into the MySQL database.
[0124] The schematic diagram of the database storage content is as Figure 2 shown. The content information stored in the database is specifically: "id" represents the item number information, "name" is the item name information, "location" is the area name where the item is located, "time" is the time information when the item is detected, "state" is the state information of the item, and the state of the item is divided into three states: visible, invisible, and blocked. "landmark" is the name of the known reference object, and "lm_state" is the position state information between the target item and the reference item, which is divided into two states: near the reference object and covered by the reference object, and can further accurately locate the position of the target item.
[0125] S38. Repeat steps S2 and S31 - S37 to achieve real - time update of the item location information in the database, so as to achieve real - time detection of the environment and real - time update of the item location information. That is, once it is found that the location of an item in the environment has changed, update the location information of the item according to the previous steps.
[0126] Embodiment 5
[0127] Step 4 is introduced in detail, specifically including:
[0128] S41. Open a separate thread to continuously monitor the user password information;
[0129] S42. When the query password is monitored, identify it as a text instruction;
[0130] S43. Pass the instruction text to the database query program through serial communication. The database query program queries the location information of the target item in the database according to the instruction requirements;
[0131] S44. Pass the query result to the voice broadcast module to broadcast the location information to the user in the form of voice.
[0132] The said voice broadcast module is the Pyttsx3 text - to - speech conversion library, which can convert any text information into voice broadcast, convenient and concise.
[0133] The voice recognition in step 4 uses the LD3320 voice recognition module.
[0134] Embodiment 6
[0135] The improved network structure of the YOLOv8 object detection model is introduced in detail, as Figure 3 shown, and it specifically includes:
[0136] Integrate the MSDA attention mechanism (multi - scale dilation attention mechanism). It captures richer context information by applying dilated convolutions at different scales and dynamically adjusts the importance of features through the multi - scale dilation attention mechanism, thereby improving the accuracy and robustness of the model in tasks such as object detection. The structural schematic diagram of the multi - scale dilation attention mechanism is as Figure 4 shown.
[0137] Introduce the EfficientNetV1 backbone network to replace the original backbone structure of YOLOv8. This backbone network integrates the advantages of depth, width, and resolution, significantly improves the inference speed of the model through a unified scaling strategy, and enhances the feature extraction ability to better identify various targets, including the detection of small - target items. The structural schematic diagram of the backbone network is as Figure 5As shown in the figure; where (a) in the figure represents the baseline network, and (b) shows the innovation of EfficientNetV1, namely the compound scaling method, which uniformly scales all three dimensions (i.e., width, depth, resolution) of the network simultaneously using a fixed ratio.
[0138] To verify the progress of the improved YOLOv8 model, a comparative experiment was conducted. The dataset used in this invention is based on the publicly available GSO indoor object dataset. Some eligible data images were selected from it, and then supplemented and improved through methods such as web crawling, on-site shooting, and data augmentation. The Indoor2024 home small object dataset containing 1232 images of 10 common household items was constructed. The dataset was divided into a training set and a validation set according to a ratio of 7:2.
[0139] The experimental environment is as follows: Windows 10 operating system, Intel CPU E5-2686 v4 processor, 60G of memory, and an NVIDIA RTX A4000 graphics card with 16G of video memory. The development tools are PyCharm, the development language version is Python 3.8, the CUDA version is 11.3, the cuDNN version is 8.0, and the Pytorch version is 1.10. The parameter configuration during the experimental training is shown in Table 1.
[0140] Table 1: Training Parameter Configuration Table
[0141]
[0142]
[0143] This experiment will comprehensively evaluate the performance differences of each model based on the following key indicators: mAP (mean Average Precision), which reflects the average precision of the model across multiple classes. Among them, mAP50 specifically refers to the average precision at a 50% IoU (Intersection over Union) threshold. mAP50-95, which is a more stringent evaluation criterion, calculates the average precision within the IoU threshold range of 50% to 95% and takes the average value to more accurately evaluate the performance of the model at different IoU thresholds. FPS (Frames Per Second), which is used to measure the processing speed of the model on specific hardware, specifically referring to the number of pictures that can be processed per second. The calculation formulas are shown as Formulas (4) and (5) below:
[0144]
[0145] Among them, m is the number of categories, and APi is the AP value of the i-th category. In particular, mAP50 represents the mAP value at the 50% IoU threshold, while mAP50-95 represents the average of the mAP values within the IoU threshold range from 50% to 95%.
[0146]
[0147] Among them, Num represents the number of frame images processed within a fixed time, and Time represents the elapsed time.
[0148] On the basis structure of the YOLOv8 object detection model, different attention mechanisms are introduced, and other structural parameters remain unchanged. Using the above evaluation indicators, comparative experiments are conducted on each attention mechanism. The experimental results are shown in Table 2, the comparison result table of attention mechanisms:
[0149] Table 2: Comparison result table of attention mechanisms
[0150]
[0151] From the comparison of the training results, it can be seen that the YOLOv8 object detection model with the added attention mechanism has changed to varying degrees. As can be seen from Table 2, when the CBAM attention mechanism is embedded, the average detection accuracy has no significant change compared with the original YOLOv8 model, mAP50 has decreased by 0.2 percentage points, and the frame rate FPS has decreased by 5; similarly, when the GAM attention mechanism is embedded, its average accuracy also decreases, but the frame rate has increased; when the MSDA attention mechanism is embedded, the model has made obvious progress, mAP50 has increased by 1.9%, mAP50-95 has increased by 1.7%, and the frame rate has decreased slightly. Therefore, it can be seen that YOLOv8n-MSDA performs best in terms of effect.
[0152] Furthermore, in order to verify the progressiveness of the improved YOLOv8 model of the present invention, the following tests are carried out on each improved part one by one under the condition of ensuring that all environmental parameters remain unchanged to prove the effectiveness of the model improvement. First, the EfficientNetV1 backbone network is replaced on the basis of the basic YOLOv8 model, and then the MSDA attention mechanism is added on this basis. The experimental results are shown in Table 3.
[0153] Table 3: Ablation experiment result table
[0154]
[0155] Table 3 conducts an experimental analysis on three groups of models from three indicators. From the comparison between Group 1 and Group 2, it can be seen that when the original YOLOv8 model is replaced with the EfficientNetV1 backbone network, the average precision of the model has been greatly improved. The mAP50 has increased by 2.8%, and the mAP50-95 has increased by 1.7%. The frame rate FPS has decreased slightly to 111. Comparing Group 2 with Group 3, Group 3 introduced the MSDA attention mechanism with the best comprehensive performance in the previous comparative experiments on the basis of Group 2. The experimental results have also been further improved compared with Group 2. The mAP50 has increased by 2.1%, and the mAP50-95 has increased by 1.4%. Similarly, there is a slight decrease in the detection speed, and the FPS is 96, but this speed does not affect the detection task of the model in the actual environment at all. The above experiments fully demonstrate that the detection model in the present invention has greatly improved the average detection precision under the condition of slightly sacrificing the detection speed, which strongly proves the rationality and advancement of the model improvement in this article.
[0156] To further reflect the performance advantages of the improved model in the present invention, this algorithm is compared with other mainstream detection models YOLOv5 and YOLOv8. The same training parameters are used in the experiments, and the experimental results are shown in Table 4. "Ours" in Table 4 is the improved detection model in this article.
[0157] Table 4: Comparison Results of Different Models
[0158]
[0159] It can be seen from the comparison results that the improved model in the present invention has the most obvious improvement in detection accuracy. Although the detection speed is lower than that of other models, it fully meets the actual application of this system. Through comprehensive analysis, this improved model takes into account both detection accuracy and detection speed and can well complete the detection and positioning tasks of indoor items.
[0160] The improved object detection model proposed in this article is based on the YOLOv8 network model. By replacing the backbone network with EfficientNetV1, it can integrate the advantages of depth, width, and resolution, significantly improve the inference speed of the model through a unified scaling strategy, and enhance the feature extraction ability to better identify various targets, including the detection of small target items. At the same time, the MSDA (Multi-Scale Dilated Attention) attention mechanism is added to apply dilated convolutions at different scales to capture richer context information and dynamically adjust the importance of features through the attention mechanism, thereby improving the accuracy and robustness of the model in tasks such as object detection. The experimental results fully show that the improved model can fully meet the requirements of actual indoor item detection and positioning. The detection accuracy has been improved to 82.6%, which is 4.9% higher than that of the original network model, and the detection effect has been significantly improved.
[0161] As Figure 6 shown, it is the detailed working flowchart of the system. All the above-mentioned content is consistent with what is shown in this flowchart. As an exact reproduction of the working process of the system of the present invention, it provides an important reference basis for understanding and analyzing the operation mechanism of the system.
[0162] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. A method for intelligently identifying and locating small items in a home scene, characterized in that: The following steps are involved: Step 1: Based on the division of the indoor detection area, frame images within a period of time in the indoor scene are collected and preprocessed; The indoor detection area division is performed when it is used for the first time, specifically: the key detection areas of the indoor environment are divided and the names of the corresponding areas are marked to obtain the area information divided by the user; the area information divided by the user includes the coordinates of the area and its corresponding name; Step 2: Send the preprocessed image to the improved YOLOv8 target detection model, detect each object and generate its detailed location information through the object location association description method, and predict the position of the occluded objects to obtain the object location information; Enter the item location information into the database for management; Step 3: Monitor user instructions in real time. After recognizing the user instructions, query the location information of the corresponding item in the database and broadcast it to the user through voice.
2. According to claim 1, a method for intelligently identifying and locating small items in a home scene is characterized in that: Divide the key detection areas of the indoor environment and mark the names of the corresponding areas to obtain the area information divided by the user, specifically: Capture the first frame of the indoor environment as the background and display it on the operation interface; The user manually selects the detection area and names the area; Automatically store the area information divided by users.
3. According to claim 1, a method for intelligently identifying and locating small items in a home scene is characterized in that: In step 1, the frame images of the indoor scene within a period of time are collected by the camera, specifically: The camera monitors the indoor scene in real time; Adjacent frame images are compared using the frame difference method. When a large change in the scene is detected that can affect the change of the object's position, frame capture of the image during this period begins.
4. According to claim 3, a method for intelligently identifying and locating small items in a home scene is characterized in that: Adjacent frame images are compared using the frame difference method. The specific processing process is as follows: Convert the frame image to grayscale, the expression is: Gray(x,y)=0.299×R(x,y)+0.587×G(x,y)+0.114×B(x,y) (1) Among them, R(x,y), G(x,y) and B(x,y) are the values of the red, green and blue color channels of the color image at the pixel point (x,y), and Gray(x,y) is the grayscale value of the corresponding grayscale image at the pixel point (x,y); the frame image is a color image; Calculate the grayscale image difference between adjacent frames. For each pixel (x, y), the difference calculation expression is: GrayDiff(x,y)=∣Gray1(x,y)-Gray2(x,y)∣ (2) Among them, Gray1(x,y) and Gray2(x,y) are the gray values of the two images at the pixel point (x,y), and GrayDiff(x,y) is the difference value of the corresponding difference image at the pixel point (x,y); When the difference value is greater than a preset threshold, the current frame is captured.
5. According to claim 1, a method for intelligently identifying and locating small items in a home scene is characterized in that: In step 1, the frame image is preprocessed, specifically: Based on the area information divided by the user obtained in step 1, the collected indoor scene frame image is cut according to the coordinate position of each area; Name each small image obtained after cutting according to the name corresponding to the area and store it.
6. The method for intelligently identifying and locating small items in a home scene according to claim 5, characterized in that: In step 2, the preprocessed frame image is sent to the improved YOLOv8 target detection model to detect each object and generate its detailed location information through the object location association description method, specifically: All the cut small images are detected using the improved YOLOv8 target detection model; The detected object in each small image is associated with the area corresponding to the small image and the reference object in the area, and the position of the object is described as its relative position to the reference object in the area; The method for describing the association of item positions is as follows: The preset area location and fixed reference objects are associated with the target object to describe the actual location of the target object in reality.
7. The method for intelligently identifying and locating small items in a home scene according to claim 6, characterized in that: The improved YOLOv8 target detection model is as follows: the MSDA attention mechanism is integrated into the YOLOv8 model, richer context information is captured by applying dilated convolutions at different scales, and the importance of features is dynamically adjusted by the MSDA attention mechanism; The EfficientNetV1 backbone network is introduced to replace the original YOLOv8 backbone structure.
8. The method for intelligently identifying and locating small items in a home scene according to claim 1, characterized in that: In step 2, the position of the occluded object is predicted, specifically: By comparing the detection results of frames within a period of time, when an object is found to have disappeared from the scene, it is preliminarily considered to be covered or taken out of the current scene; Combined with the position coordinates of the disappeared object in the previous frame, the Kalman filter is used to predict its coordinate position after disappearance to determine its physical position in reality. The Kalman filter state equation is: x_k = A ×x_(k-1) + B ×u_(k-1) + w_(k-1) (3) Among them, x_k is the state vector at time k, A is the state transfer matrix, x_(k-1) is the state vector at time k-1, B is the control input matrix, u_(k-1) is the control input at time k-1, and w_(k-1) is the process noise.
9. The method for intelligently identifying and locating small items in a home scene according to claim 1, characterized in that: In step 2, the item location information is entered into the database for management, specifically: Item location information includes item name, location area, detection time, and whether it is covered; Enter the item location information into the MySQL database, and repeat steps 2 and 3 to achieve real-time update of the item location information in the database; Step 3 is as follows: Loop monitoring of user password information, when monitoring the user password, identify it as a command text; The command text is transmitted to the database query program through serial communication, and the database query program queries the location information of the target object in the database according to the command requirements; The query results are transmitted to the user in the form of voice to report the location information.
10. A small item intelligent identification and positioning system in a home scene, characterized in that: include: The area division module is used to divide the key detection areas of the indoor environment and mark the names of the corresponding areas when used for the first time, and obtain the area information divided by the user; the area information divided by the user includes the coordinates of the area and its corresponding name; An image preprocessing module is used to preprocess the frame images collected within a period of time in the indoor scene; The object detection module is used to send the preprocessed image into the improved YOLOv8 object detection model, detect each object and generate its detailed location information through the object location association description method, and predict the position of the occluded objects to obtain the object location information; A location information database is used to enter the location information of items into a database for management; A speech recognition module is used to recognize the real-time monitoring user commands; A database query module is used to query the location information of the items stored in the location information database after recognizing the user instruction, and query the location information of the corresponding items in the database; The voice broadcast module is used to broadcast the location information of the corresponding items queried from the database query module to the user through voice.