A room state detection method and system based on image recognition

By tracking the motion trajectory of core objects and performing refined image analysis, the problem of inaccurate detection in complex environments by existing image recognition methods has been solved, achieving more accurate room status detection and reducing false alarms and missed alarms.

CN120612737BActive Publication Date: 2025-10-17CHENGDU SHANMEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511121069.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-17
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing image recognition methods struggle to accurately distinguish between changes in the actual state of an object and image changes caused by factors such as lighting, surface defects, and dust in complex indoor environments, leading to false alarms and missed detections in room status detection.

Method used

By acquiring image sequences inside the room, identifying and tracking the movement trajectories of preset core objects, and combining image analysis with the final state to be detected, the repositioning action of the core objects is verified, and the state of non-core objects or environmental areas is detected to generate the final state report of the room.

Benefits of technology

It significantly improves the accuracy and reliability of room status detection, reduces false alarms and missed alarms, and improves the efficiency of automated management and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612737B_ABST
    Figure CN120612737B_ABST
Patent Text Reader

Abstract

The application discloses a room state detection method and system based on image recognition, and belongs to the field of image recognition detection, which comprises the following steps: acquiring an image sequence of the inside of a room, identifying a preset core object from the image sequence, and tracking the motion trail of the preset core object; determining that the room enters a final detection state based on the stationary state of the preset core object, the closed state of the door, and the state of no person existing in the room; backtracking the motion trail to verify whether the preset core object completes a preset homing action; analyzing the image of the room in the final detection state to detect the state of non-core objects or environmental areas in the room; and generating a final state report of the room according to the homing action verification result and the environmental area state detection result. The application can effectively distinguish the image changes caused by the real state change of the object from the image changes caused by the environmental factors, significantly improve the accuracy and reliability of the room state detection, and reduce the false positives and false negatives.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image recognition detection, in particular to a room state detection method and system based on image recognition. BACKGROUND

[0002] In a specific indoor environment, image recognition technology is often used to automatically detect the state of objects in the space. However, in practical applications, inaccurate recognition is often caused by complex physical environmental factors. These factors may include indoor lighting conditions, surface characteristics of the detected objects, micro changes in the objects during use, and interference from small particles in the environment, all of which can cause deviations between image information and the true physical state, affecting the reliability of the detection.

[0003] Taking the piano room state monitoring of a certain music school as an example, dozens of independent piano rooms are used at a high frequency. To facilitate management, an automatic detection system based on image recognition is used to verify whether the room has returned to the standard state after each user leaves. It is usually necessary to confirm whether the piano cover is completely closed, whether the music stand is folded and placed in the designated area of the corner, whether the piano stool is pushed back under the piano, whether the indoor lighting is turned off, and whether the floor is left with any objects. The system usually uses a fixed camera installed in the corner of the room ceiling to capture a panoramic image of the room after the user leaves and closes the door, and compares it with the pre-stored standard state template image to complete the detection.

[0004] However, such automatic detection systems often report false positives. After excluding the possibility of camera hardware failure and software program execution error, there are still false positives. The underlying reason may be that different objects in the room have different surface materials, different reflection characteristics of light, or unevenness of light, resulting in a large difference in brightness information in the original data collected by the image sensor, reducing the accuracy of image comparison. With frequent use of the piano room, fingerprints, sweat stains left on the piano varnish surface, or fine traces left by the music stand on the metal stand. These surface imperfections, under certain lighting angles, change the local light scattering or absorption characteristics. The existing image recognition method is a pixel or feature level matching between the image and a static, idealized template. This method cannot establish a causal relationship between the physical world, and cannot distinguish between a visual change on the image caused by a real change in the state of the object (such as the piano cover being opened) or a series of complex physical environmental factors such as reflection of light on the fingerprint, random dust obstruction. These situations can all lead to inaccurate detection.

[0005] In view of the above problems, the existing technology needs to be improved. SUMMARY

[0006] In order to solve the problem of inaccurate detection of the state of objects in a room in the prior art, the present application provides a room state detection method and system based on image recognition, which can effectively distinguish between image changes caused by changes in the true state of objects and image changes caused by complex environmental factors such as light, surface flaws, and dust, significantly improving the accuracy and reliability of room state detection, reducing false positives and false negatives, and thus improving the efficiency of automated management and user experience.

[0007] The present application is realized by the following technical solutions:

[0008] The present application provides a room state detection method based on image recognition, comprising the following steps: acquiring an image sequence of the interior of a room, which records the movement process of a preset core object in the room; identifying the preset core object from the image sequence and tracking the movement trajectory of the preset core object; determining that the room enters a final state to be detected, based on the stationary state of the preset core object, the closed state of the door, and the absence of personnel in the room; backtracking the movement trajectory to verify whether the preset core object has completed the preset homing action; analyzing the image of the room in the final state to be detected to detect the state of non-core objects or environmental areas in the room; and generating a final state report of the room based on the homing action verification result and the environmental area state detection result.

[0009] Further, the step of backtracking the movement trajectory to verify whether the preset core object has completed the preset homing action comprises: identifying the preset core object, assigning an identity to the preset core object based on its visual features; associating the movement trajectory with the identity of the preset core object; determining whether the identity of the preset core object that performed the homing action is consistent with the identity of the target preset core object; and if the identity of the preset core object that performed the homing action is not consistent with the identity of the target preset core object, determining that the homing action has not been completed.

[0010] Further, the step of analyzing the image of the room in the final state to be detected to detect the state of non-core objects or environmental areas in the room comprises: dividing the image of the room in the final state to be detected into regions to obtain a preset detection region, which includes the surface region of the core object; analyzing the local light variation characteristics and texture characteristics of each surface region of the core object; identifying abnormal regions in the local light variation characteristics or texture characteristics that do not match the preset reference of the surface of the core object based on the local light variation characteristics and texture characteristics; and determining whether the abnormal regions have an independent contour or a specific size to identify the residual objects on the surface of the core object.

[0011] Further, after the step of performing, for each surface region of the core object, an analysis of the local lighting variation characteristic and the texture characteristic of the region, the method further comprises: determining a preset reference update condition of the region, the preset reference update condition comprising a preset periodic time interval or a continuous detection period during which no residual is identified on the region; and when the preset reference update condition is met, taking the analyzed local lighting variation characteristic and the texture characteristic as the updated preset reference of the region.

[0012] Further, the step of determining whether the abnormal region has an independent contour or a specific size to identify the residual on the surface of the core object comprises: performing a light transmission or reflection characteristic analysis on the abnormal region to identify whether the region causes a local change in background light transmission or reflection; performing an analysis on the local lighting distribution around the abnormal region to identify whether the region forms a fixed shadow region; and determining whether the abnormal region has an independent contour or a specific size, or the abnormal region causes a local change in background light transmission or reflection, or the abnormal region forms a fixed shadow region, to identify the residual on the surface of the core object.

[0013] Further, the step of dividing the room image in the final to-be-detected state into regions to obtain a preset detection region, the preset detection region comprising the surface region of the core object, comprises: identifying feature points or structural edges of the core object in the room image in the final to-be-detected state; calculating the position and pose of the core object in the image according to the feature points or structural edges; determining an initial corresponding range of the surface region of the core object in the image based on the position and pose of the core object and in combination with a preset geometric structure of the core object; analyzing the image content in the initial corresponding range to determine whether there is a local defect or abnormality in the surface region that does not conform to the preset integrity of the core object; and adjusting the initial corresponding range according to the determination result to obtain the preset detection region.

[0014] Further, the step of identifying the abnormal region that does not conform to the preset reference of the surface of the core object in the local lighting variation characteristic or the texture characteristic comprises: calculating a first difference degree between the local lighting variation characteristic and the preset reference of the surface of the core object; calculating a second difference degree between the texture characteristic and the preset reference of the surface of the core object; and identifying the abnormal region according to a relationship between the first difference degree and a preset first threshold value or according to a relationship between the second difference degree and a preset second threshold value.

[0015] Further, the step of calculating the first difference degree between the local illumination change characteristic and the preset reference of the surface of the core object comprises: performing mesh division on the surface area of the core object to obtain a plurality of mesh areas; for each mesh area in the plurality of mesh areas, extracting statistical features of the local illumination change characteristic and the preset reference, the statistical features including brightness mean and brightness variance; calculating a difference value between the statistical features of the local illumination change characteristic and the corresponding statistical features of the preset reference according to the extracted statistical features of each mesh area; and performing weighted aggregation on the difference values of all mesh areas to obtain the first difference degree.

[0016] Further, the step of identifying the preset core object, assigning an identity of the preset core object according to visual features of the preset core object comprises: identifying structural feature points and texture areas of the preset core object; extracting geometric relationship features, local texture features and spectral reflection features of the preset core object according to the structural feature points and the texture areas; fusing the geometric relationship features, the local texture features and the spectral reflection features to form comprehensive identity features of the preset core object; comparing the comprehensive identity features with a preset object identity feature library, and determining the identity of the preset core object in combination with historical identity information of the preset core object; and updating the corresponding comprehensive identity features in the object identity feature library in a case that the identity of the preset core object is determined and there is no residual object on the surface of the object.

[0017] The second aspect of the present application also provides a room state detection system based on image recognition, which is used to execute the room state detection method based on image recognition, and the system comprises: an acquisition module configured to acquire an image sequence of a room interior, the image sequence recording a movement process of a preset core object in the room; an identification and tracking module configured to identify the preset core object from the image sequence and track a movement trajectory of the preset core object; a state determination module configured to determine that the room enters a final detection state, the determination being based on a stationary state of the preset core object, a closed state of a door of the room and a state that no person exists in the room; a homing verification module configured to backtrack and analyze the movement trajectory to verify whether the preset core object completes a preset homing action; a region detection module configured to analyze a room image in the final detection state to detect states of non-core objects or environmental regions in the room; and a report generation module configured to generate a final state report of the room according to a homing action verification result and an environmental region state detection result.

[0018] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0019] The present application scheme can effectively distinguish between image changes caused by real state changes of objects and image changes caused by environmental factors (such as illumination, surface flaws and dust), significantly improve the accuracy and reliability of room state detection, and reduce false positives and false negatives.

[0020] The present application effectively distinguishes the relationship between image changes and physical facts by introducing backtracking analysis of the core object motion trajectory, accurate determination of the final state to be detected in the room, and fine detection of the state of non-core objects and environmental areas, thereby effectively distinguishing between image changes caused by changes in the real state of objects and image changes caused by complex environmental factors such as light, surface flaws, and dust, significantly improving the accuracy and reliability of room state detection, reducing false positives and false negatives, and improving the efficiency of automated management and user experience. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor. In the drawings:

[0022] Figure 1 A flowchart of a room state detection method based on image recognition provided by the present application.

[0023] Figure 2 A structure diagram of a room state detection system based on image recognition provided by the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present application clearer, the following will further describe the present application in combination with embodiments and drawings. The exemplary embodiments of the present application and their descriptions are only used to explain the present application and do not limit the present application.

[0025] The image recognition method in the prior art has difficulty in effectively distinguishing between image visual changes caused by changes in the real state of objects and image visual illusions caused by complex physical environmental factors (such as differences in the reflection of light by object materials, subtle flaws on the surface of objects, local shadows produced by fixed light sources, and airborne particulate matter) when automatically detecting the state of indoor objects. This results in frequent false positives or false negatives when determining whether the room has returned to the preset standard state, thereby affecting the reliability of automated detection.

[0026] In a certain scenario, an image recognition-based automatic detection system is deployed in a high-frequency music training facility to verify the state of the piano room after the user leaves. The room image is captured by a fixed camera and compared with the pre-stored standard template. Due to the diversity of the materials in the piano room (the piano surface is high-gloss paint, the music stand is matte metal, and the piano stool is velvet fabric, which have different reflection characteristics to light). Under the illumination of the fixed overhead light, the piano surface forms a high-light area, while the piano stool surface hardly reflects light, resulting in uneven image brightness information. In addition, the surface of the object may change slightly during use, such as fingerprints or sweat stains on the piano paint, which can change the local light scattering under certain lighting conditions, causing a closed piano lid to appear as a bright spot in the image due to fingerprint reflection, which is misinterpreted as an open piano lid. At the same time, the fixed light source will form fixed shadows at the folding joints of the music stand or between the piano keys, causing the structure of the objects in the shadow to be blurred and unable to be clearly identified, thus determining that they are not in place. Tiny dust particles suspended in the air drift through the image capture moment, causing temporary shading or diffraction of light, forming random noise points on the image. These noise points are similar in image features to real litter on the ground, causing a clean floor to be misjudged as having litter. These phenomena collectively result in the inability to accurately distinguish between visual changes in the image that are due to changes in the true state of the object or environmental physical factors.

[0027] Based on the above problems, the embodiment of the present application proposes a room state detection method combining motion trajectory analysis and final state image analysis to more comprehensively and accurately evaluate the room state. As shown in Figure 1 , the method specifically includes the following steps:

[0028] 1. Obtain an image sequence of the interior of the room, which records the motion process of the pre-set core objects in the room;

[0029] The pre-set core object refers to a specific object in the room that needs to be focused on and its state needs to be detected. It can be furniture, equipment, or any object with a clear requirement for returning to its original position, such as a piano, a music stand, a piano stool, etc. Its main purpose is to accurately manage and verify the state of key objects in the room.

[0030] 2. Identify the pre-set core object from the image sequence and track the motion trajectory of the pre-set core object;

[0031] A series of image frames obtained continuously within a period of time are used as the image sequence. The camera continuously records the activities in the room to record the dynamic change process of the pre-set core object in the room, providing complete time dimension data for subsequent motion trajectory tracking and behavior analysis.

[0032] A motion trajectory refers to the set of continuous positions of a preset core object in an image sequence over time. It can be represented by a center point coordinate sequence, a bounding box position sequence, or a key point displacement sequence. Its main purpose is to record and describe the movement path of the preset core object from its initial position to its final position, providing a basis for determining whether it has completed a specific action.

[0033] 3. Determine whether the room has entered the final state to be inspected, based on the static state of the preset core objects, the closed state of the door, and the absence of people in the room;

[0034] The final state to be detected refers to the state in which the room enters a stable condition suitable for final state detection after the user activity is completed. The judgment is based on the static state of the preset core objects, the closed state of the door, and the presence of no people in the room. For example, when all core objects in the room stop moving, the door is in the closed position, and no human activity is detected in the room, it is mainly to ensure that the environment is stable and undisturbed when performing the final state analysis of the room, so as to avoid misjudgment due to continuous movement or the presence of people.

[0035] 4. Backtrack and analyze the motion trajectory to verify whether the preset core object has completed the preset homing action;

[0036] The homing action refers to the preset core objects being moved to their preset, standardized storage position or posture. It can be achieved by pushing the items back to their original position, folding them for storage, or closing the covers. For example, the piano stool is pushed back under the piano and the music stand is folded and placed in the corner. It is mainly to verify whether the preset core objects are restored to the designated position in accordance with management requirements to ensure the cleanliness and order of the room.

[0037] 5. Analyze the final image of the room to be inspected and detect the status of non-core objects or environmental areas in the room;

[0038] The status of non-core objects or environmental areas refers to the conditions of other objects or specific areas in the room besides the preset core objects. This can be achieved by detecting whether there are any leftovers on the ground, whether the lights are turned off, whether the windows are closed, etc. Its main purpose is to conduct a comprehensive assessment of the overall environment of the room, make up for the omissions that may exist when focusing only on core objects, and ensure the overall compliance of the room.

[0039] 6. Generate the final status report of the room based on the homing action verification results and the environmental area status detection results.

[0040] The embodiments of the present application continuously acquire image sequences of the interior of the room, which record in detail the movement process of the preset core objects in the room. Based on these image sequences, the preset core objects can be identified and their movement trajectories can be accurately tracked. This provides historical data support for subsequent homing action verification, which can not only determine the final position of the object, but also understand the path and method of reaching the position. When the room enters the final detection state, i.e., the preset core objects are in a static state, the door is closed, and no one is present in the room, deep analysis is started. At this time, the movement trajectory recorded before is analyzed in retrospect to verify whether the preset core objects have completed the preset homing action. The retrospective analysis can effectively distinguish whether the object is actively homing or passively in a certain position, thereby avoiding the possible misjudgment of relying solely on the final image. At the same time, the room image in the final detection state is also analyzed to detect the state of non-core objects or environmental areas in the room. This step is a supplement to the homing verification of core objects, ensuring a comprehensive assessment of the overall environment of the room, such as detecting whether there are any left-over objects on the floor, whether the lights are off, etc. Finally, the verification result of the homing action is combined with the state detection result of the non-core objects or environmental areas to obtain a comprehensive final state detection result of the room. This method of combining dynamic process analysis and static environment detection can more accurately judge the actual physical state of the room, effectively avoiding the misjudgment of the single image comparison method under the interference of complex lighting, surface defects or environmental particles, thereby improving the accuracy and reliability of the detection. Through multi-stage and multi-dimensional data analysis, accurate detection of the state of the room is achieved.

[0041] In a specific scenario, such as a music training room, the preset core objects include a piano, a piano stool, and a music stand. First, a wide-angle camera installed in the corner of the room continuously acquires image sequences inside the room, and records the complete process of the user moving the piano stool, folding the music stand, and other operations in the room at a predetermined frequency. When the user leaves the room, the image sequences are analyzed in real time to identify the piano, piano stool, and music stand, and a target tracking algorithm is used to continuously record their position and attitude changes in the room coordinate system, thereby generating their respective motion trajectories. Subsequently, the room state is continuously monitored. Once it is detected that the piano stool and the music stand have both stopped moving, the door is closed by the door magnetic sensor feedback, and there is no person present in the room by the human presence sensor, it is determined that the room enters the final detection state. After entering the final detection state, the homing verification process is immediately started. The motion trajectories of the piano stool and the music stand are analyzed. For example, for the piano stool, check if its trajectory contains a continuous path from in front of the piano to the area below the piano, and the final position is within the preset homing area below the piano. For the music stand, analyze whether its trajectory contains a path from the use position to the designated storage area in the corner of the wall, and verify whether its final attitude is in the folded state. At the same time, the room image in the current final detection state is analyzed in detail. This includes image processing of the piano surface, piano stool surface, music stand surface, and ground area to detect whether there are non-core objects or left objects. Analyze whether there are abnormal light spots or texture changes on the piano cover surface to determine whether there are left objects or fingerprints; detect whether there are foreign objects inconsistent with the background on the ground area. Finally, the homing verification results of the piano stool and the music stand are integrated with the detection results of the non-core objects and the environment area in the room to obtain a comprehensive room state detection result.

[0042] The above scheme of the embodiment of the present application can effectively solve the misjudgment problem caused by light changes, object surface flaws, or environmental particulate matter interference in the existing image recognition method in a complex indoor environment. By introducing the backtracking analysis of the motion trajectory of the preset core object, it can be accurately judged whether the object has completed the preset homing action, avoiding the misjudgment that may be caused by relying only on the final static image.

[0043] By analyzing the motion trajectory, it is verified whether the preset core object completes the preset homing action. Specifically, by identifying each preset core object, assigning a unique identity based on its visual features, and associating the motion trajectory with the identity, it is determined whether the identity of the preset core object performing the homing action is consistent with the identity of the target preset core object, which can ensure that the homing action is completed by the correct object. However, in fact, in the case of multiple similar preset core objects, it is possible to mistakenly associate the motion trajectory of one preset core object with the identity of another preset core object, resulting in a false homing verification result. In order to avoid misjudgment due to object confusion or identity error, in some possible implementations, the step of analyzing the motion trajectory to verify whether the preset core object completes the preset homing action includes:

[0044] a. Identify the preset core object, and assign an identity to the preset core object based on the visual features of the preset core object. The visual features of the preset core object refer to the inherent visual properties of the object itself that can be obtained by an image sensor.

[0045] b. Associate the motion trajectory with the identity of the preset core object.

[0046] c. Determine whether the identity of the preset core object performing the homing action is consistent with the identity of the target preset core object. The target preset core object refers to a specific preset core object that is expected to perform the homing action in a specific scenario, which is to clarify the subject of the homing action and ensure the accuracy of the verification.

[0047] d. If the identity of the preset core object performing the homing action is not consistent with the identity of the target preset core object, it is determined that the homing action is not completed.

[0048] When identifying the preset core object, a deep learning model such as a convolutional neural network (CNN) is used to detect and classify each object in the image. For each identified preset core object, its visual features (such as color histogram, SIFT feature point, SURF feature point, ORB feature point, local binary pattern (LBP) texture feature, or three-dimensional point cloud data) are extracted. Based on these visual features, a unique identity is assigned to each preset core object. The position change of each preset core object with an identity in the image sequence is continuously tracked to determine which preset core object's motion trajectory conforms to the preset homing action path. Once the object performing the homing action is identified, its assigned identity during the motion process is queried, and this identity is directly compared with the identity of the preset target preset core object. If the identity of the preset core object performing the homing action is not consistent with the identity of the target preset core object, it is determined that the homing action is not completed; if it is consistent, the homing action is completed.

[0049] The overall analysis of the room image is easily affected by factors such as uneven lighting, shadows, differences in the reflective properties of the object surface, and interference from tiny particles, which can lead to the inability to accurately identify the residue on the surface of the core object, thereby affecting the accuracy and reliability of the room status detection. In some possible implementations, the steps of analyzing the room image in the final state to be detected and detecting the status of non-core objects or environmental areas in the room include:

[0050] a. The final room image to be inspected is divided into predefined inspection areas, which include the surface of the core object. Using image processing technology, the predefined inspection area is a collection of pixels representing the surface of the core object in the image. This limits the inspection scope to the surface of the core object, eliminating interference from background and other non-critical areas, thereby improving the specificity and efficiency of the inspection.

[0051] b. Analyze the local illumination variation and texture characteristics of each surface area of ​​the core object. Local illumination variation refers to image features that reflect the absorption, reflection, and scattering of light by the object's surface, capturing the visual differences caused by changes in lighting conditions.

[0052] c. Based on local illumination and texture characteristics, identify abnormal areas where these characteristics do not conform to the preset benchmarks of the core object surface. The preset benchmarks for the core object surface refer to the illumination and texture characteristics that the core object should possess in a normal, clean, or standard state. These can be pre-collected or learned, providing a reference standard to determine whether the current state is abnormal.

[0053] d. Determine whether the abnormal area has a distinct outline or specific size to identify artifacts on the core object's surface. An abnormal area is defined as a region where local illumination variation or texture characteristics significantly differ from a preset baseline. An independent outline or specific size means that the abnormal area exhibits a discernible boundary in the image and its size falls within the preset artifact size range, distinguishing true artifacts from artifacts caused by uneven lighting, shadows, or surface imperfections.

[0054] For the surface area of the core object, analyzing its local illumination variation characteristics and texture characteristics can capture surface abnormalities that are difficult to detect with the naked eye, such as stains, scratches, or tiny particulate matter. Local illumination variation characteristics reflect the absorption, reflection, and scattering of light by the object surface, while texture characteristics describe the roughness, smoothness, and pattern distribution of the object surface. Comparing the analyzed local illumination variation characteristics and texture characteristics with a pre-set reference can identify abnormal areas that differ from the normal state. The pre-set reference can be the illumination and texture characteristics of the core object in a clean state, or the stable state characteristics formed after a period of use. By judging the contours and sizes of the abnormal areas, it is possible to distinguish between real residues and artifacts caused by uneven lighting, shadow blocking, or surface imperfections. Specifically, the room image in the final state to be detected can be regionally divided to obtain a pre-set detection area, which includes the surface area of the core object. A deep learning model, such as a semantic segmentation model based on a convolutional neural network, is used to classify the room image at the pixel level, identifying the precise location and boundaries of the core object (such as a piano, music stand, or piano stool) in the image. From these identified core object areas, the surface areas are further extracted. For example, for a piano, the lid, key area, and top of the body can be identified as surfaces. For each surface area of the core object, the local illumination variation characteristics and texture characteristics of the area are analyzed. The local brightness histogram, mean and variance of each surface area are calculated as local illumination variation characteristics. At the same time, the gray level co-occurrence matrix (GLCM) or local binary pattern (LBP) features are extracted as texture characteristics. The response of the surface to light and its microscopic structure are quantified. According to the local illumination variation characteristics and texture characteristics, abnormal areas that do not match the pre-set reference of the core object surface are identified in the local illumination variation characteristics or texture characteristics. The current analyzed local illumination and texture characteristics are compared with the pre-stored reference characteristics representing the clean state of the core object. If the difference between the mean brightness or texture characteristics of a certain area and the reference value exceeds a pre-set threshold, the area is marked as an abnormal area. The abnormal area is judged for its independent contour or specific size to identify residues on the surface of the core object. The identified abnormal areas are subjected to connected component analysis to extract the contour information and area of each connected component. If an abnormal area has a clear independent contour and its area or aspect ratio falls within the pre-set residue size range (e.g., greater than a certain minimum number of pixels and less than a certain maximum number of pixels), it can be determined that the abnormal area is a residue on the surface of the core object. The embodiments of the present invention focus on the surface area of the core object by performing fine regional division on the room image in the final state to be detected, thereby effectively excluding the interference of the background and non-critical areas.

[0055] As the surface state of the core object changes over time, such as changes in lighting conditions, wear and tear due to long-term use, etc., these changes will cause the original preset reference to be no longer applicable, thereby affecting the accuracy of the residual detection. Therefore, how to update the preset reference of the surface area of the core object according to the actual situation to adapt to environmental changes is a technical problem to be solved. Therefore, in one possible embodiment, the steps of updating the preset reference of the surface area of the core object include:

[0056] a. Determine the preset reference update condition of the area, and the preset reference update condition includes a preset periodic time interval or no residual is identified in the area in a continuous detection period. The preset reference update condition refers to a specific condition for triggering the update of the preset reference of the surface area of the core object. It is realized in a time-driven or event-driven manner, which ensures that the preset reference can adapt to changes in the environment or the state of the object. The preset periodic time interval refers to automatically triggering the preset reference update according to a fixed time period, which is used to cope with slowly changing environmental factors, such as seasonal changes in lighting or natural aging of the object. No residual is identified in the area in a continuous detection period refers to that in multiple continuous detection periods, the specific area is not judged to have residual, indicating that the area is in a stable and clean state, which is used to use the actual cleaning state of the object as the basis for updating the preset reference, avoiding inaccurate reference update when there is residual.

[0057] b. When the preset reference update condition is met, the local lighting change characteristics and texture characteristics analyzed are used as the updated preset reference of the area. The local lighting change characteristics refer to the local characteristics of the lighting intensity, direction or distribution in the surface area of the core object, which reflect the reflection or absorption of the object surface to light. The texture characteristics refer to the visual pattern or structural characteristics of the surface area of the core object. The updated preset reference refers to the new reference data of the surface state of the core object for subsequent residual detection after the update operation. The local lighting change characteristics and texture characteristics analyzed are used as new reference values, which can accurately judge based on the latest state of the object.

[0058] After analyzing the local lighting variation characteristics and texture characteristics of each surface region of the core object, further determine the preset reference update conditions of the region. These conditions can include preset periodic time intervals, such as automatically triggering a reference update every fixed time, which can adapt to seasonal changes in ambient light or slow aging of the object. In addition, the update condition can also include that no residual is identified in the region during the continuous detection period, which means that when a region remains clean in continuous detection, the current state is considered to be the stable normal state of the region, thereby triggering the reference update. Once these preset update conditions are met, the current analysis of the local lighting variation characteristics and texture characteristics will be used as the updated preset reference of the region. In actual scenarios, if the piano surface appears slight wear or gloss changes due to long-term use, the new preset reference will take it into account to avoid misjudging these normal changes as residual. After the preset reference of the surface region of the core object is dynamically updated, the subsequent residual identification step (i.e., identifying abnormal regions according to the difference between the local lighting variation characteristics or texture characteristics and the preset reference, and judging whether the abnormal region has an independent contour or a specific size to identify the residual) will be based on a more accurate and more adaptive reference to the current environment. This can effectively distinguish between visual illusions caused by environmental changes or natural aging of the object and real residuals, thereby reducing false positives and improving the reliability and adaptability of room state detection. This dynamic updating mechanism enables the entire room state detection method to continuously self-calibrate and optimize, ensuring accurate detection results in complex and changing environments. The embodiment of the present application can dynamically update the preset reference of the core object surface state according to the actual changes of the core object surface state. This solves the problem that the original preset reference is no longer applicable due to changes in the core object surface state over time, which affects the accuracy of residual detection. Through periodic updating or updating based on a clean state, it can continuously adapt to environmental light changes, object wear, and other factors, ensuring that residual detection is always based on the latest and accurate reference state, thereby improving the accuracy and long-term reliability of residual detection.

[0059] The abnormal region identified by the lighting and texture characteristics may not accurately distinguish between the residual on the surface of the core object and the characteristics of the object itself, such as stains, wear, or manufacturing defects on the surface of the object, which may be misjudged as a residual, resulting in false detection results. In one possible embodiment, the step of determining whether the abnormal region has an independent contour or a specific size to identify the residual on the surface of the core object includes:

[0060] a. Analyzing the light transmission or reflection characteristics of the abnormal region to identify whether the region causes local changes in the background light transmission or reflection. Light transmission or reflection characteristic analysis refers to evaluating the optical interaction between the abnormal region and the surrounding environment by measuring the degree of absorption, scattering, or reflection of incident light by the abnormal region, to distinguish the effects of different materials or forms of objects on light propagation. Local changes in the background light transmission or reflection refer to the presence of the abnormal region causing detectable changes in the intensity, direction, or spectral composition of the background light behind or around it, revealing the physical intervention of the abnormal region on the light path.

[0061] b. Analyzing the local illumination distribution around the abnormal region to identify whether the region forms a fixed shadow area. Local illumination distribution analysis refers to evaluating the intensity, gradient, and direction of illumination in the abnormal region and its immediate vicinity to identify shadow features formed by objects blocking light. Fixed shadow area refers to the dark area formed around the abnormal region under the illumination of a specific light source, with stable shape, position, and intensity, indicating the three-dimensional structure of the abnormal region and its blocking effect on light.

[0062] c. Judging whether the abnormal region has an independent contour or specific size, or whether the abnormal region causes local changes in the background light transmission or reflection, or whether the abnormal region forms a fixed shadow area, to identify the residual on the surface of the core object. Independent contour or specific size refers to the abnormal region presenting a separable and identifiable boundary from the surrounding background in the image, and its area, perimeter, or aspect ratio conforming to the preset size range, distinguishing the residual from the subtle defects of the object itself in terms of geometric morphology.

[0063] By introducing multi-dimensional feature analysis, the preliminary identified abnormal regions are further differentiated. The light transmission or reflection characteristics of the abnormal regions are analyzed to identify whether the region causes local changes in the transmission or reflection of background light. A transparent glass cup or a translucent plastic bag will change the transmission of background light, while a metal coin or a highly glossy paper sheet will enhance local reflection. Misjudgments caused by surface gloss or color changes of the object itself can be excluded, as the characteristics of the object itself usually do not cause such significant local changes in the background light. The local light distribution around the abnormal region is analyzed to identify whether the region forms a fixed shadow area. As three-dimensional objects, residues usually cast shadows under light. By analyzing the shape, location, and stability of these shadows, it can be further confirmed whether the abnormal region is an actually existing residue. Unlike the inherent shadows formed by the structure of the object itself (e.g., piano key gaps), the shadows produced by residues often have irregularity or dynamics. The above light transmission or reflection characteristics, the formation of fixed shadow areas, and whether the abnormal region has an independent contour or a specific size are comprehensively judged. The logical OR relationship is adopted, meaning that as long as any one of the conditions is met, the abnormal region can be identified as a residue on the surface of the core object. This multi-feature fusion judgment method compensates for the limitations of single-feature judgment. Based on the previously identified local light change characteristic and texture characteristic abnormal regions, the present scheme introduces judgment of the optical characteristics and shadow features of the abnormal regions, thereby more accurately distinguishing true residues from inherent characteristics of the object itself (such as stains, wear, manufacturing defects) or environmental interference (such as dust, fingerprints). For an image region that has been preliminarily identified as abnormal, the original pixel data collected by the image sensor can be used to analyze the light transmission or reflection characteristics. A reference background region can be set, and the differences in average brightness value, color channel distribution, or highlight reflection intensity of the pixels in the abnormal region and the reference background region can be compared. If the average brightness value of the abnormal region is significantly higher than the surrounding background, and it exhibits a mirror reflection characteristic, it indicates that the background light reflection has been locally changed, which may indicate a high-reflectivity residue. Conversely, if the brightness value of the abnormal region is significantly lower than the surrounding background, and it exhibits a blurred or dull transmission effect, it may indicate a transparent or translucent residue that has caused a local change in the transmission of background light. The local light distribution around the abnormal region is analyzed to identify whether the region forms a fixed shadow area. This is achieved by gradient analysis or edge detection on the abnormal region and its adjacent pixels. The brightness gradient of the pixels within a certain range outside the boundary of the abnormal region is calculated. If there is a significant and continuous brightness drop area, and the shape and direction of the area are consistent with the preset light source direction, it can be determined that the abnormal region forms a fixed shadow. A deep learning model can be used to train it to identify different shapes and intensities of shadow patterns, thereby more robustly judging the presence of shadows.The abnormal region is comprehensively judged whether it has an independent contour or a specific size, or the abnormal region causes a local change of background light transmission or reflection, or the abnormal region forms a fixed shadow area, to identify the residual on the surface of the core object. The abnormal region is binarized and analyzed by a connected domain, the contour thereof is extracted and the area and aspect ratio thereof are calculated. If the area of the region is within a preset effective residual size range and the aspect ratio thereof is within a reasonable range, the region has a specific size and an independent contour. The geometric features are logically ORed with the aforementioned light transmission / reflection characteristic analysis results and shadow area analysis results. Even if a residual is not obvious in contour due to its material, as long as it causes a local change of background light or forms a detectable shadow, it can still be identified as a residual. This multiple discrimination mechanism can effectively improve the identification rate of various residuals and reduce false positives.

[0064] The region division of the room image cannot guarantee that the divided region can accurately cover the surface area of the core object, especially when the core object has a posture change or is partially blocked, which can easily cause deviation in subsequent analysis and detection, such as dividing a region that does not belong to the surface of the core object, introducing interference information, or missing part of the surface area of the core object, resulting in failure to detect a residual in the region. In one possible embodiment, a step of dividing the room image in the final state to be detected to obtain a preset detection region, the preset detection region including the surface area of the core object, is proposed.

[0065] a. In the room image in the final state to be detected, feature points or structural edges of the core object are identified. The feature points or structural edges refer to local pixel patterns or geometric boundaries in the image that have unique and repeatable detection, providing stable reference information of the core object in the image for accurate positioning and posture estimation.

[0066] b. The position and posture of the core object in the image are calculated according to the feature points or structural edges.

[0067] c. Based on the position and posture of the core object, and in combination with the preset geometric structure of the core object, the initial corresponding range of the surface area of the core object in the image is determined. The preset geometric structure refers to the inherent shape, size, and relative position relationship between parts of the core object in three-dimensional space, which provides a basis for mapping the two-dimensional projection of the core object in the image to its real three-dimensional surface, thereby accurately determining the surface area thereof.

[0068] d. Analyzing the image content in the initial corresponding range to determine whether the surface area has local defects or abnormalities that do not conform to the preset integrity of the core object. The preset integrity refers to the continuous, undamaged, and unobstructed physical state of the surface of the core object in an ideal or standard state, which serves as a basis for determining whether the surface area of the core object is accurate and whether it contains abnormalities, thereby ensuring the effectiveness of the detection area.

[0069] e. Adjusting the initial corresponding range according to the determination result to obtain a preset detection area.

[0070] In the room image in the final state to be detected, feature points or structural edges of the core object are identified. These feature points or structural edges are inherent to the object and remain relatively stable under different perspectives, laying the foundation for subsequent accurate calculations. Based on these identified feature points or structural edges, the position and pose of the core object in the image are calculated. Associating two-dimensional image information with the state of the object in three-dimensional space enables accurate grasp of the actual orientation and spatial position of the object. In combination with the pre-set geometry of the core object, the initial corresponding range of the surface area of the core object in the image is determined. The pre-set geometry provides accurate three-dimensional morphological information of the object. By combining the calculated position and pose with this three-dimensional morphology, the surface of the object can be accurately projected onto the image plane, obtaining a preliminary, close-to-real surface area range. To further improve accuracy, the image content in this initial corresponding range is analyzed to determine whether there is a local defect or anomaly in the surface area that does not conform to the pre-set integrity of the core object. Adaptive verification of the initial division result can discover deviations in the initial range caused by lighting, occlusion, or minor changes in the object itself. If part of the object surface is occluded, or there is distortion during image acquisition, these incomplete or abnormal situations can be identified. According to the judgment result, the initial corresponding range is adjusted to obtain the final preset detection area. This ensures that the detection area closely fits the actual visible surface of the core object, avoiding the inclusion of non-object surface areas in the detection range and preventing the omission of areas on the object surface that need to be detected. The preset detection area obtained at this time can more accurately cover the surface of the core object. Feature extraction algorithms such as SIFT (Scale-Invariant Feature Transform) or SURF (Speeded-Up Robust Features) can be used to identify feature points such as corner points, edge points, or texture spots of the core object (e.g., a piano) in the image. At the same time, methods such as Canny edge detection or Hough transform are used to extract structural edges of the piano lid, body, or stool. In combination with PnP (Perspective-n-Point) algorithms or Iterative Closest Point (ICP) algorithms, the three-dimensional position and pose of the piano in the current image are calculated, including its translation and rotation information in space. Based on the calculated position and pose of the piano, and in combination with the pre-stored three-dimensional geometric model of the piano (e.g., an accurate CAD model or three-dimensional point cloud data), the three-dimensional model is projected onto the current image plane, thereby determining the initial corresponding range of the piano surface area in the image, such as a polygonal region composed of multiple vertices. Subsequently, the image content in this initial corresponding range is analyzed. Image segmentation techniques or deep learning models are used to determine whether there is a local defect in the region that does not conform to the pre-set integrity of the piano, such as a missing corner of the piano lid, or the presence of abnormal shadows, highlights, or foreign matter occlusions that are not part of the piano surface.The presence of significant differences can also be detected by comparing the actual image within the initial corresponding range with a preset image template of the piano surface in the state without residual objects. According to the analysis result, the initial corresponding range can be adjusted. For example, if it is determined that a certain region is blocked, the detection range of the region can be reduced; if it is found that the initial range fails to completely cover a visible surface of the piano, the range can be expanded. Such adjustment can be achieved by morphological operations (such as dilation and erosion) or algorithms based on boundary fitting, and finally an accurate preset detection region is obtained for subsequent residual object detection. The embodiment effectively avoids the detection deviation that can be introduced by the traditional simple region division method in the case of changes in the posture of the object or partial blocking. Through a series of refined steps, accurate division of the surface region of the core object is achieved, thereby overcoming the limitations of the traditional simple region division method in the case of changes in the posture of the object or partial blocking.

[0071] Due to the influence of factors such as uneven lighting, differences in object surface materials, and slight physical changes, directly comparing the local lighting variation characteristics or texture characteristics with the preset reference can be easily disturbed by noise, resulting in reduced accuracy of residual object identification. In one possible embodiment, the steps of identifying abnormal regions in the local lighting variation characteristics or texture characteristics that do not match the preset reference of the core object surface include:

[0072] a. Calculate the first difference degree between the local lighting variation characteristics and the preset reference of the core object surface. The local lighting variation characteristics refer to the brightness, contrast, light and shadow distribution, etc. of a specific region in the image, which are related to the visual attributes of the lighting, and capture the visual differences caused by changes in the lighting conditions or surface state of the object. The preset reference of the core object surface refers to the reference standard of the local lighting variation characteristics and texture characteristics of the core object surface region in the normal state without residual objects, which can be image feature data collected and stored during initialization, or average feature data updated after multiple anomaly-free detections, providing a reference for comparison and judgment. The first difference degree refers to a quantitative indicator of the deviation between the currently detected local lighting variation characteristics and the preset reference of the core object surface, which objectively measures the abnormal degree of lighting variation.

[0073] b. Calculate the second difference degree between the texture characteristics and the preset reference of the core object surface. The second difference degree refers to a quantitative indicator of the deviation between the currently detected texture characteristics and the preset reference of the core object surface, which is used to objectively measure the abnormal degree of texture variation.

[0074] c. identifying the abnormal region according to the relationship between the first difference and a preset first threshold value, or according to the relationship between the second difference and a preset second threshold value. The abnormal region refers to a region on the surface of the core object where the local illumination variation characteristic or the texture characteristic significantly deviates from the preset reference, which is manifested as brightness abnormality, shadow, highlight point, or texture structure change, etc., indicating a potential location where there may be residual or surface defect.

[0075] A first difference degree between the current local illumination variation characteristic and the preset reference of the surface of the core object is calculated. The deviation degree of the illumination feature is converted into a quantifiable numerical value, so that the subsequent judgment has objective basis. A second difference degree between the current texture characteristic and the preset reference of the surface of the core object is also calculated, to quantify the deviation degree of the texture structure. Since the difference degrees of the illumination and the texture in two different dimensions are quantified, the possible abnormalities of the surface of the object can be more comprehensively captured. According to the relationship between the calculated first difference degree and the preset first threshold value, or the relationship between the second difference degree and the preset second threshold value, the abnormal area is identified. Based on the threshold value judgment mechanism, non-substantial changes caused by environmental noise, slight illumination fluctuation or inherent differences in material can be effectively filtered out, so as to focus on the significant abnormalities indicating the left-over. The problem that the traditional direct comparison method is easily disturbed is avoided, and the accuracy of identifying the abnormal area is significantly improved. Assuming that the core object is a piano, the surface area of the piano has been identified and isolated through image processing technology. First, the current image data of the surface area of the piano is obtained. In order to calculate the first difference degree between the local illumination variation characteristic and the preset reference of the surface of the piano, local brightness analysis can be performed on the surface area, for example, the brightness value of each pixel point or small block area is calculated, and compared with the pre-stored piano surface brightness reference collected under standard illumination conditions. The sum of the absolute values of the pixel-level brightness difference, or the sum of squares of the regional average brightness difference is used to represent, so as to obtain the first difference degree. At the same time, in order to calculate the second difference degree between the texture characteristic and the preset reference of the surface of the piano, texture feature extraction can be performed on the surface area, and the texture feature vector is extracted using the gray level co-occurrence matrix or the local binary pattern algorithm, and compared with the pre-stored piano surface texture reference collected in the state without left-over. The Euclidean distance or cosine similarity between the feature vectors can be used to represent the comparison, so as to obtain the second difference degree. According to the relationship between the calculated first difference degree and the preset first threshold value, or the relationship between the second difference degree and the preset second threshold value, the abnormal area is identified. If the first difference degree exceeds the preset first threshold value, it indicates that the illumination of the area has a significant abnormality, which may be caused by fingerprints, water stains, etc.; or if the second difference degree exceeds the preset second threshold value, it indicates that the texture structure of the area has changed significantly, which may be caused by paper scraps, coins and other left-overs. When either condition is met, the area is marked as an abnormal area, and further judgment is performed to identify the left-over. The method can objectively quantify the deviation degrees of the illumination and the texture by calculating the first difference degree between the local illumination variation characteristic and the preset reference of the surface of the core object, and the second difference degree between the texture characteristic and the preset reference of the surface of the core object.According to the relationship between the difference degree and the preset threshold, the abnormal area can be identified, and non-substantial interference caused by uneven light, difference in surface material of the object, and small physical changes, etc. can be effectively filtered out, thereby significantly improving the accuracy of identifying the abnormal area which is inconsistent with the preset reference of the surface of the core object in the local light change characteristic or the texture characteristic.

[0076] By judging the overall difference degree, the accuracy of the abnormal area identification can be affected by uneven light, shadow blocking, etc. leading to inaccurate difference degree calculation results. In one possible embodiment, the step of calculating the first difference degree between the local light change characteristic and the preset reference of the surface of the core object comprises:

[0077] a. Dividing the surface area of the core object into a plurality of grid areas.

[0078] b. For each grid area in the plurality of grid areas, extracting statistical features of the local light change characteristic and the preset reference, the statistical features including brightness mean and brightness variance.

[0079] c. According to the statistical features extracted in each grid area, calculating the difference value between the statistical features of the local light change characteristic and the corresponding statistical features of the preset reference. The difference value refers to the quantitative deviation degree between the statistical features of the local light change characteristic and the corresponding statistical features of the preset reference, which can accurately measure the deviation degree of the current local light from the standard state.

[0080] d. Weighted aggregation of the difference values of all grid areas to obtain the first difference degree. Weighted aggregation refers to combining the difference values of the plurality of grid areas according to the preset weight to form a comprehensive evaluation index, which comprehensively considers the influence of different areas on the overall abnormal judgment, so that the final first difference degree can more accurately reflect the overall abnormal situation of the surface of the core object.

[0081] In each independent grid region, the local illumination variation characteristics and the statistical features of the preset reference are extracted, including the brightness mean and the brightness variance. The brightness mean provides the average illumination intensity information of the local region, while the brightness variance reveals the uniformity or volatility of the illumination in the region. By considering both statistical features, the local illumination situation can be more comprehensively described. A region may have normal overall brightness but local highlights or shadows, which will be reflected in the variance. The extraction of multi-dimensional statistical features makes the description of local illumination more robust, reducing the errors caused by single pixel values or simple average values. According to the statistical features extracted in each grid region, the difference value between the statistical features of the local illumination variation characteristics and the corresponding statistical features of the preset reference is calculated, quantifying the deviation between the current illumination state of each small region and the ideal reference. Since the difference value is calculated at the local grid level, it can accurately reflect the abnormalities of a specific small region, for example, a fingerprint may only affect the brightness mean and brightness variance of a few grid regions, without significantly affecting the global statistics of the entire surface region. The difference values of all grid regions are aggregated by weighting to obtain the first difference degree. The introduction of weighted aggregation allows different weights to be assigned according to the importance of different grid regions or their impact on the final judgment. For example, the key regions of the core object (such as the center part of the piano lid) can be assigned higher weights, while the edge or unimportant regions can be assigned lower weights. This weighting mechanism enables the final first difference degree to more accurately reflect the overall abnormality of the core object surface, while effectively suppressing false positives caused by non-critical regions or random noise. Through the synergistic effect of the above grid division, local statistical feature extraction, difference value calculation, and weighted aggregation, the problem of inaccurate difference degree calculation caused by uneven illumination, shadow blocking, or small surface defects in traditional methods can be effectively overcome. For the surface of the piano lid, the corresponding region in the image can be extracted. The surface region is divided into grids, and for each grid region, the statistical features of the local illumination variation characteristics and the preset reference are extracted, including the brightness mean and the brightness variance. The brightness mean and the brightness variance of all pixels in the grid region can be calculated. The brightness mean reference and the brightness variance reference of the corresponding grid region are obtained from the preset reference data. Then, according to the current statistical features (brightness mean and brightness variance) extracted in each grid region and the corresponding statistical features of the preset reference, the difference value between them is calculated. The absolute difference between the current brightness mean and the reference brightness mean, as well as the absolute difference between the current brightness variance and the reference brightness variance, are calculated, and the weighted sum of the two absolute differences is obtained to obtain the difference value of the grid region. The difference values calculated for all grid regions are aggregated by weighting to obtain the final first difference degree. In the weighted aggregation, weights can be assigned according to the importance of the grid regions on the surface of the core object.For the grid of the central region of the piano cover, a higher weight can be assigned because these regions are more prone to fingerprints or residues; while for the edge or inconspicuous regions, a lower weight can be assigned. By multiplying the difference value of each grid region by its corresponding weight, and adding up all the weighted difference values, the final first difference degree can be obtained. The first difference degree value can then be compared with a preset first threshold to determine whether there is an abnormal region on the surface of the core object. Through the calculation of the local difference value and the weighted aggregation, the embodiment of the present application comprehensively evaluates the abnormality of the entire surface, and reduces the influence of random noise and non-critical regions on the detection result.

[0082] Identity identification relying on visual features is easily disturbed by factors such as light changes, occlusions, deformations, etc., leading to identity identification errors, and further affecting the verification of subsequent homing actions, especially in the complex environment of the piano room, a single visual feature may not be enough to distinguish different core objects, or in the case of stains, wear and tear on the surface of the object, the visual features will change, leading to recognition errors. In one possible embodiment, a step of identifying a preset core object and assigning an identity identification of the preset core object according to visual features of the preset core object is proposed, which includes:

[0083] a. Identifying the structural feature points and texture regions of the preset core object.

[0084] b. Extracting geometric relationship features, local texture features and spectral reflection features of the preset core object according to the structural feature points and texture regions.

[0085] c. Fusing the geometric relationship features, local texture features and spectral reflection features to form a comprehensive identity feature of the preset core object.

[0086] d. Comparing the comprehensive identity feature with a preset object identity feature library, and combining the historical identity information of the preset core object to determine the identity identification of the preset core object.

[0087] e. Updating the corresponding comprehensive identity feature in the object identity feature library when the identity identification of the preset core object is determined and there is no residue on the surface of the object.

[0088] The structural feature points refer to key points of the object in the image that have stability and repeatability; the texture region refers to an image region of the object surface that has a specific repetitive pattern or detail distribution; the geometric relationship feature refers to the relative position, distance, angle, or topological relationship between the structural feature points in space; the local texture feature refers to the statistical characteristics of the pixel gray scale or color distribution in the texture region; the spectral reflection feature refers to the absorption and reflection characteristics of the object surface to different wavelengths of light; the comprehensive identity feature refers to a composite feature vector formed by fusing multiple independent features, which can fully represent the identity of the object, and can be constructed by using feature-level fusion, decision-level fusion, or deep learning feature fusion; the object identity feature library refers to a feature dataset that is pre-stored and used to compare and identify the identities of various preset core objects, and contains the comprehensive identity features of each known object and the corresponding identity labels; the historical identity information refers to the identity label of a certain preset core object determined in a previous detection or identification process and the related confidence or timestamp information.

[0089] The structural feature points and texture regions of the preset core object are identified, and geometric relationship features, local texture features and spectral reflection features are extracted. These features extracted from different dimensions are fused to form a comprehensive identity feature of the preset core object. The fusion mechanism makes up for the shortcomings of a single feature by the advantages of other features, thereby constructing an identity representation that is resistant to environmental interference. When the local texture features are damaged due to light, the geometric relationship features and the spectral reflection features can still provide reliable identification basis. The comprehensive identity feature is compared with the preset object identity feature library, and the historical identity information of the preset core object is combined to determine the identity of the object. Through comparison with the feature library, the comprehensive feature of the current object can be matched with the feature template of the known object. The introduction of historical identity information enables auxiliary judgment by using the past identification records of the object when facing ambiguous or similar identification results, thereby improving the accuracy and continuity of identification and avoiding identity jumping caused by instantaneous environmental changes. In the case that the identity of the preset core object is determined and there is no residual on the surface of the object, the corresponding comprehensive identity feature in the object identity feature library is updated. The object identity feature library can be self-adapted to the actual use of the object and the long-term changes of the environment. When the object surface appears slight wear or stains, and the visual features change slightly, the update without residual interference can learn and adapt to these changes, ensuring that the identity of the object can still be accurately identified even if the appearance of the object changes slightly. The identification errors caused by changes in the appearance of the object in the prior art are avoided, thereby ensuring the reliability of subsequent homing action verification. In a complex environment such as a music room, the object materials are various, the light is fixed, and the surface micro changes are easily affected. A single visual feature is easy to fail. The present embodiment can capture the essential properties of the object from multiple dimensions by combining geometric, texture and spectral information, distinguish different objects, and resist interference such as light, occlusion, deformation, and surface stains and wear. The ability to dynamically update the object identity feature library can adapt to long-term changes of the object.

[0090] By identifying the structural feature points and texture regions, and extracting the geometric relationship features, local texture features and spectral reflection features, the inherent properties and surface details of the object are captured from multiple dimensions. These multi-dimensional features are fused to form a comprehensive identity feature that is resistant to interference such as light changes, occlusion, deformation, and surface stains and wear. The combination of comparison with the object identity feature library and the assistance of historical identity information improves the accuracy and robustness of the identity of the preset core object. In addition, in the case that the identity is determined and there is no residual on the surface of the object, the object identity feature library is dynamically updated to adapt to the long-term changes of the object, ensuring the continuous accuracy of identity recognition. This provides reliable object identity information for subsequent homing action verification, thereby improving the overall reliability of room state detection and reducing the occurrence of false positives.

[0091] Based on the same inventive concept of the above method embodiments, the present embodiments propose an image recognition-based room state detection system for implementing the above method embodiments, as shown in Figure 2 The system comprises:

[0092] An acquisition module is configured to acquire an image sequence of the interior of the room, the image sequence recording a movement process of a preset core object in the room.

[0093] An identification and tracking module is configured to identify the preset core object from the image sequence and track a movement trajectory of the preset core object.

[0094] A state determination module is configured to determine that the room enters a final detection state, the determination being based on a stationary state of the preset core object, a closed state of a door of the room, and a state of no person existing in the room.

[0095] A homing verification module is configured to backtrack the movement trajectory and verify whether the preset core object completes a preset homing action.

[0096] A region detection module is configured to analyze a room image in the final detection state and detect a state of a non-core object or an environmental region in the room.

[0097] A report generation module is configured to generate a final state report of the room according to a homing action verification result and an environmental region state detection result.

[0098] The above detailed description further describes the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only a specific implementation of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A room status detection method based on image recognition, characterized in that: The following steps are involved: Acquire an image sequence of the interior of the room, wherein the image sequence records a movement process of a preset core object in the room; identifying the preset core object from the image sequence and tracking the motion trajectory of the preset core object; Determining that the room enters a final state to be detected, the determination being based on the static state of the preset core object, the closed state of the door, and the absence of any person in the room; Backtracking and analyzing the motion trajectory to verify whether the preset core object has completed the preset homing action; Analyze the room image in the final state to be inspected to detect the state of non-core objects or environmental areas in the room; generating a final status report of the room according to the homing action verification result and the environmental area status detection result; The step of backtracking and analyzing the motion trajectory to verify whether the preset core object has completed the preset homing action includes: Identifying the preset core object, and assigning an identity tag to the preset core object based on visual features of the preset core object; Associating the motion trajectory with the identity of the preset core object; Determining whether the identity identifier of the preset core object performing the homing action is consistent with the identity identifier of the target preset core object; If the identity identifier of the preset core object performing the homing action is inconsistent with the identity identifier of the target preset core object, it is determined that the homing action is not completed; The step of analyzing the room image in the final state to be detected and detecting the state of non-core objects or environmental areas in the room includes: Divide the room image in the final state to be detected into regions to obtain a preset detection region, where the preset detection region includes a surface region of a core object; For each surface area of ​​the core object, analyzing the local illumination variation characteristics and texture characteristics of the area; identifying, based on the local illumination variation characteristics and the texture characteristics, an abnormal area in the local illumination variation characteristics or the texture characteristics that does not conform to a preset reference on the surface of the core object; Determine whether the abnormal area has an independent outline or a specific size to identify the remains on the surface of the core object.

2. The room status detection method based on image recognition according to claim 1, characterized in that: After performing the step of analyzing the local illumination variation characteristics and texture characteristics of each surface area of ​​the core object, the method further includes: Determining a preset benchmark update condition for the area, the preset benchmark update condition including that no remains are identified in the area within a preset periodic time interval or a continuous detection period; When the preset benchmark update condition is met, the local illumination change characteristics and texture characteristics obtained by the analysis are used as the updated preset benchmark for the region.

3. The room status detection method based on image recognition according to claim 1, characterized in that: The step of determining whether the abnormal area has an independent outline or a specific size to identify the remains on the surface of the core object includes: Analyzing the light transmission or reflection characteristics of the abnormal area to identify whether the area causes a local change in background light transmission or reflection; Analyzing the local illumination distribution around the abnormal area to identify whether the area forms a fixed shadow area; Determine whether the abnormal area has an independent outline or a specific size, or whether the abnormal area causes a local change in the transmission or reflection of background light, or whether the abnormal area forms a fixed shadow area, so as to identify the remains on the surface of the core object.

4. The room status detection method based on image recognition according to claim 1, characterized in that: The step of dividing the room image in the final state to be detected into regions to obtain a preset detection region, wherein the preset detection region includes the surface region of the core object, comprises: identifying feature points or structural edges of the core object in the final room image in the to-be-detected state; Calculating the position and posture of the core object in the image according to the feature points or structure edges; Based on the position and posture of the core object and in combination with a preset geometric structure of the core object, determining an initial corresponding range of the surface area of ​​the core object in the image; Analyzing the image content within the initial corresponding range to determine whether the surface area has a local defect or anomaly that is inconsistent with the preset integrity of the core object; According to the judgment result, the initial corresponding range is adjusted to obtain the preset detection area.

5. The room status detection method based on image recognition according to claim 1, characterized in that: The step of identifying an abnormal area in the local illumination change characteristic or the texture characteristic that does not conform to a preset reference on the surface of the core object comprises: Calculating a first difference between the local illumination variation characteristic and a preset reference on the surface of the core object; calculating a second difference between the texture characteristic and a predetermined reference on the surface of the core object; The abnormal area is identified based on a relationship between the first difference and a preset first threshold, or based on a relationship between the second difference and a preset second threshold.

6. The room status detection method based on image recognition according to claim 5, characterized in that: The step of calculating a first difference between the local illumination variation characteristic and a preset reference on the surface of the core object comprises: Meshing the surface area of ​​the core object to obtain a plurality of mesh areas; For each of the plurality of grid areas, extracting statistical features of the local illumination variation characteristic and the preset benchmark, wherein the statistical features include a brightness mean and a brightness variance; Calculating, based on the statistical features extracted from each grid area, a difference between the statistical features of the local illumination variation characteristics and corresponding statistical features of the preset benchmark; The difference values ​​of all grid areas are weightedly aggregated to obtain the first difference degree.

7. The room status detection method based on image recognition according to claim 1, characterized in that: The step of identifying the preset core object and assigning an identity tag to the preset core object according to the visual features of the preset core object includes: Identifying structural feature points and texture areas of the preset core object; Extracting geometric relationship features, local texture features, and spectral reflectance features of the preset core object based on the structural feature points and the texture area; fusing the geometric relationship feature, the local texture feature, and the spectral reflectance feature to form a comprehensive identity feature of the preset core object; Comparing the comprehensive identity feature with a preset object identity feature library and combining the historical identity information of the preset core object to determine the identity identifier of the preset core object; When the identity identifier of the preset core object is determined and there is no residue on the surface of the preset core object, the corresponding comprehensive identity feature in the object identity feature library is updated.

8. A room status detection system based on image recognition, based on the room status detection method based on image recognition according to any one of claims 1 to 7, characterized in that: The system includes: An acquisition module, configured to acquire an image sequence of the interior of a room, wherein the image sequence records a movement process of a preset core object in the room; an identification and tracking module, configured to identify the preset core object from the image sequence and track the motion trajectory of the preset core object; A state determination module is used to determine whether the room has entered a final state to be detected, the determination being based on whether the preset core object is in a stationary state, the door is closed, and no one is present in the room; A homing verification module, configured to retrospectively analyze the motion trajectory and verify whether the preset core object has completed the preset homing action; An area detection module is used to analyze the room image in the final state to be detected and detect the state of non-core objects or environmental areas in the room; A report generation module is used to generate a final status report of the room based on the homing action verification result and the environmental area status detection result.

Citation Information

Patent Citations

  • Table tennis motion analysis method

    CN114862900A

  • Double-background modeling remnant detection method based on multiple backtracking verification

    CN115909199A