Image recognition-based control method and system, vehicle, and storage medium
By acquiring the status and image information of the display interface, extracting feature information, and controlling according to control commands, the problem of poor user experience in image recognition is solved, and efficient and intelligent image recognition control is achieved.
Patent Information
- Application Number
- CN202010474388.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-05-29
AI Technical Summary
In existing technologies, the user experience of image recognition is not perfect.
By acquiring the status information of the display interface and judging changes, image information is acquired and feature information is extracted. The feature information is then controlled according to control instructions, including voice instructions and server instructions.
It improves the user experience of image recognition, reduces processor computing power consumption and energy consumption, and enables image recognition control without software adaptation.
Smart Images

Figure CN113741769B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of application control technology, and in particular to a control method based on image recognition, a computer-readable storage medium, a control system based on image recognition, and a vehicle. Background Art
[0002] In recent years, with the continuous development of technology, the application of image recognition in life has become increasingly widespread. However, the inventors found that the user experience of image recognition in the existing technology is not perfect. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, one object of the present invention is to propose a control method based on image recognition, which aims to solve the shortcomings of the prior art to a certain extent.
[0004] A second objective of the present invention is to provide a computer-readable storage medium.
[0005] A third object of the present invention is to provide a control system based on image recognition.
[0006] A fourth object of the present invention is to provide a vehicle.
[0007] In order to solve the above problems, the control method based on image recognition of the first aspect of the embodiment of the present invention includes: obtaining status information of the display interface and making a judgment; if the status information changes, obtaining image information of the display interface and extracting feature information of the image information; obtaining control instructions for the feature information; and controlling the feature information according to the control instructions.
[0008] According to the image recognition-based control method provided by an embodiment of the present invention, the control instructions for feature information can automatically obtain image information of the display interface based on changes in the display interface state information, and then extract feature information from the image information. After obtaining the control instructions for the feature information, the feature information is controlled according to the control instructions. This improves the user experience of image recognition to a certain extent.
[0009] In some embodiments, obtaining the image information of the display interface and extracting the feature information of the image information includes obtaining the image information of the display interface and comparing it with the previously obtained image information, identifying the image area that has changed in the image information, extracting the feature information of the image area and updating it to the previously extracted feature information.
[0010] In some embodiments, acquiring image information of the display interface and extracting feature information of the image information includes recording a video of the display interface and extracting feature information of the video.
[0011] In some embodiments, when the state information of the display interface stops changing and the duration is greater than or equal to a first preset time, the video recording of the display interface is stopped;
[0012] The extracting feature information of the video includes extracting feature information of key frames in the video.
[0013] In some embodiments, obtaining the status information of the display interface includes obtaining display status information of the application on the display interface;
[0014] If the state information changes, acquiring the image information of the display interface and extracting the characteristic information of the image information includes: if the display state of the application on the display interface changes, and the duration of the changed display state is greater than or equal to a second preset time, starting to acquire the image information of the display interface and extracting the characteristic information of the image information, and when the duration of the changed display state is greater than or equal to a third preset time, stopping acquiring the image information of the display interface;
[0015] Wherein, the second preset time is shorter than the third preset time.
[0016] In some embodiments, if the display state of the application on the display interface changes, it includes that the application displayed on the display interface changes or the current display interface of the application changes.
[0017] In some embodiments, the feature information includes a text control button area, a graphic control button area, and a text input area;
[0018] The acquiring of the control instruction for the characteristic information includes acquiring a voice instruction, an instruction issued by a server, an instruction transmitted by a third party, or an instruction automatically generated by the system;
[0019] The controlling of the characteristic information according to the control instruction includes performing click, slide, or text input operations on the characteristic information;
[0020] After controlling the characteristic information according to the control instruction, the method further includes feeding back the control result to the user.
[0021] In some embodiments, the feature information includes coordinate position information on the display interface;
[0022] The method further includes obtaining operation information of a user when controlling the feature information, and simulating the user's operation information when controlling the feature information according to the control instruction;
[0023] The operation information includes click action information, slide action information, and text input action information.
[0024] The computer-readable storage medium of the second embodiment of the present invention stores a computer program thereon, and when the computer program is executed, it implements the control method based on image recognition described in the above embodiment.
[0025] The image recognition-based control system of the third aspect of the embodiment of the present invention includes: a status information acquisition module, used to obtain status information of the display interface and make a judgment; an image information acquisition module, used to obtain image information of the display interface and extract feature information of the image information when the status information changes; a control instruction acquisition module, used to obtain control instructions for the feature information; and a control module, used to control the feature information according to the control instructions.
[0026] In the image recognition-based control system provided by an embodiment of the present invention, the image information acquisition module automatically acquires image information of the display interface based on changes in the display interface status information acquired by the status information acquisition module, thereby extracting feature information from the image information. After the control instruction acquisition module acquires the control instruction for the feature information, the control module controls the feature information according to the control instruction. This, to a certain extent, improves the user experience of image recognition-based control instructions for feature information.
[0027] The vehicle of the fourth embodiment of the present invention includes a display device and the image recognition-based control system described in the above embodiment.
[0028] In a vehicle according to an embodiment of the present invention, employing the image recognition-based control system of the above embodiment, the image information acquisition module automatically acquires image information of the display interface based on changes in the display interface status information acquired by the status information acquisition module, thereby extracting feature information from the image information. After the control instruction acquisition module acquires control instructions for the feature information, the control module controls the feature information in accordance with the control instructions. This improves the user experience of image recognition to a certain extent.
[0029] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0031] Figure 1 is a flow chart of a control method based on image recognition according to an embodiment of the present invention;
[0032] Figure 2 This is a flow chart of a control method applied to a vehicle-mounted central control display screen according to an embodiment of the present invention;
[0033] Figure 3 is a coordinate position diagram according to an embodiment of the present invention;
[0034] Figure 4 is a schematic diagram of a control system based on image recognition according to an embodiment of the present invention;
[0035] Figure 5 is a schematic diagram of a control system applied to a vehicle-mounted central control display screen according to an embodiment of the present invention;
[0036] Figure 6 FIG. 1 is a schematic diagram of a vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0038] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.
[0039] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In distributed computing environments, program modules may be located in both local and remote computer storage media, including memory devices.
[0040] In the present invention, "module", "device", "system" and the like refer to related entities applied to a computer, such as hardware, a combination of hardware and software, software or software in execution, etc. Specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program and / or a computer. In addition, an application or script program running on a server, or a server can all be an element. One or more elements can be in an execution process and / or thread, and an element can be localized on a computer and / or distributed between two or more computers, and can be run by various computer-readable media. An element can also communicate through local and / or remote processes based on a signal having one or more data packets, for example, a signal from a data packet interacting with another element in a local system, a distributed system, and / or a signal from a network on the Internet that interacts with other systems via a signal.
[0041] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device. In the absence of further limitations, the elements defined by the phrase "include..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.
[0042] The image recognition-based control method in the embodiment of the present invention corresponds to a computer program product, which is installed on a smart terminal device and is used to realize voice control of a third-party application installed on the smart terminal device (voice control of the third-party application can be realized without the need to customize or adaptively debug or modify the third-party application). The smart terminal is equipped with a display screen or the terminal device can project a display interface for user interactive operations, such as any smart hardware such as a smart phone, tablet computer, PC, car terminal, smart home, projector, etc., and the present invention is not limited to this.
[0043] The image recognition-based control technology solution provided by the present invention can be applied to the control of display interface application software, such as controlling the opening, interface switching, application removal and other operations of music playback software, map navigation software, etc., and can also be applied to the control of display interface content, such as picture processing.
[0044] like Figure 1FIG2 is a flow chart of a control method based on image recognition according to an embodiment of the present invention. The control method based on image recognition according to the embodiment of the present invention comprises at least steps S101-S104.
[0045] Step S101: Obtain status information of the display interface and make a judgment.
[0046] Among them, the display interface of the present invention can be the display interface of a display screen, or it can be a projected display interface. The status information of the display interface can be the application interface currently displayed on the display interface. For example, the display interface of the present invention is the display interface of a vehicle-mounted central control display screen, and the status information of the display interface is that the current display interface displays an application interface, which can be a map navigation interface, a music broadcast interface, a game interface, etc., or it can be the current interface of an application, such as a music list interface, a music search interface, a music playback interface, etc. of a music player. To obtain the status information of the display interface and make a judgment, the method can be to obtain the information of the currently running application, the display information of the application on the display interface, the information of the current display interface of the application, etc. After obtaining the status information of the display interface, further analysis and judgment are performed.
[0047] Step S102: If the state information changes, image information of the display interface is acquired and feature information of the image information is extracted.
[0048] Specifically, after obtaining the status information of the display interface, analyze whether the status information has changed. For example, after obtaining the application currently displayed on the display interface or the current display interface information of the application, analyze and compare it with the previously obtained status information to determine whether the displayed application has changed or whether the display interface of the application has changed. For example, if the application displayed on the display interface switches from a navigation application to a music playback application, it can be determined that the status information has changed; or if the current display interface of the music playback application changes from a playback interface to a music search interface, it can be determined that the status information has changed. After determining that the status information has changed, obtain the image information of the display interface and extract the feature information in the image information. The image information of the display interface can be obtained by taking a screenshot of the current display interface, or by taking a screenshot of only part of the display interface. For example, the display interface may display multiple different types of information and information from different applications. For example, a vehicle's central control display can simultaneously display weather information, vehicle information (including in-car temperature, battery life, etc.), navigation information, music playback information, and multimedia information (such as WeChat and Weibo information). Some display content, such as vehicle information, is resident, and some applications are not displayed full-screen, occupying only a portion of the display interface. Therefore, when the status information of the display interface changes, such as when a new application is opened but only occupies a portion of the display interface, if the image information of the entire display interface is still obtained and feature information is extracted, redundant information processing will result, resulting in a waste of computing power. At the same time, it may cause unnecessary extension of information processing time, delaying the user's time.
[0049] The feature information of the image information is extracted, and the feature information may include text-based control buttons, graphic-based control buttons, text input areas, and the like. For example, information such as a return icon control button, previous / next page control buttons, fast forward / rewind button information, a progress bar, and a text input box is extracted from the image information. It is understood that the buttons can be circular, rectangular, or other irregularly shaped, or simply in text form. They can also be unconventional buttons, such as the commonly used left-swipe to go to the previous page, right-swipe to go to the next page, left-side up / down swipe to adjust brightness, right-side up / down swipe to adjust volume, and double-click to play / pause. In this case, the buttons of this embodiment can also be unconventional control buttons. Information about unconventional control buttons can be obtained by analyzing information such as the display coordinate area of the application or the application on the display interface, and determining the operation coordinate area of the left-side, double-click, and up-down swipe based on the user's conventional left-right swipe, double-click, and up-down swipe operations, and using the operation coordinate area and the operation as the feature information of the image information. This achieves comprehensive extraction of feature information of the image information on the display interface.
[0050] Step S103: Acquire a control instruction for the characteristic information.
[0051] In an embodiment of the present invention, a control instruction for feature information is obtained. The control instruction may be a voice control instruction from a user or a control instruction sent from a server. For example, a control instruction for feature information (click, slide, etc.) may be extracted from a voice instruction sent by a user.
[0052] Step S104: Control the characteristic information according to the control instruction.
[0053] Specifically, after obtaining the control instruction for the characteristic information, the characteristic information is controlled according to the control instruction. For example, if the user's voice control instruction is to play song B by singer A, then text input will be made into the music search box in the characteristic information and searched according to the control instruction, and the retrieved song B will be played.
[0054] The image recognition method provided by an embodiment of the present invention can automatically acquire image information of a display interface based on changes in the display interface's state information, and further extract feature information from the image information. After obtaining control instructions for the feature information, the feature information is controlled accordingly. This, to a certain extent, improves the user experience of image recognition.
[0055] For example, in the prior art, users want to control applications by issuing voice commands. However, application software requires adaptation and other modifications to implement voice control functionality. According to the technical solution provided by the embodiments of the present invention, image recognition technology is used to extract feature information from image information, and then control this feature information based on control instructions. As a result, without the need for adaptive software modifications, it is possible to perform image recognition feature information on most conventional software applications and then control them based on voice control instructions.
[0056] It is understood that some embodiments of the present invention are described using user voice control as an example. However, the application of the present invention is not limited to the field of user voice control technology. It can also be applied to a variety of fields that can replace user operations. For example, when purchasing tickets, it can help users automatically refresh, automatically enter verification codes, automatically place orders and exit, and perform a series of simulated user operations. The image recognition-based control method of the present invention uses image recognition to find feature information and perform control. It does not require customization, adaptive debugging, or modification of the application program. The feature information can be controlled according to the received control instructions, thereby enhancing the intelligent experience.
[0057] Among them, the control method based on image recognition provided by the present invention obtains the status information of the display interface, and only when the status information changes, the subsequent image information acquisition and feature information extraction process is performed, thereby avoiding repeated processing of information, reducing the consumption of processor computing power, etc., and also reducing energy consumption.
[0058] The control method based on image recognition provided by an embodiment of the present invention starts to acquire image information and extract feature information when the status information changes. For example, when a new application is starting to load or a display interface of an application is starting to load, it is detected that the status information of the display interface has changed. At this time, the image information is started to be acquired and feature information is extracted, which makes full use of the period of time when the application is starting to load or a display interface of the application is starting to load to acquire image information and extract feature information. Therefore, the image acquisition and feature extraction of the loaded content can be performed in advance.
[0059] In actual operation, extracting feature information from an image requires identifying the image and performing a large number of calculations, which consumes a lot of computing power and also requires a certain amount of computing time. It is not always possible to identify and extract the feature information. Therefore, according to the control method based on image recognition provided by an embodiment of the present invention, when it detects that a new application is being loaded or a new display interface of an application is being loaded, it starts to acquire the image and extract feature information, which can speed up the recognition response speed, greatly save the user's waiting time, and even make the user not feel the time of acquiring image information and extracting feature information, thus achieving the purpose of responding to user needs at any time. It avoids interference with the user's normal operations and improves the user experience.
[0060] In some embodiments, obtaining the image information of the display interface and extracting the feature information of the image information includes obtaining the image information of the display interface and comparing it with the image information obtained previously, identifying the image area that has changed in the image information, extracting the feature information of the image area and updating it to the feature information extracted previously.
[0061] Specifically, to improve information processing efficiency and avoid redundant information processing, an embodiment of the present invention further provides a solution, including obtaining image information of the display interface and comparing it with previously obtained image information, identifying the image region that has changed in the image information, extracting feature information of the image region, and updating it with the previously extracted feature information. In many cases, only a portion of the displayed content on the display interface has changed. In this case, only the changed portion needs to be processed, and the feature information of that portion needs to be extracted and updated with the previously extracted feature information. This greatly reduces the amount of information processing, saves data processing time, reduces user waiting time, and further improves the user experience.
[0062] In some embodiments, acquiring image information of the display interface and extracting feature information of the image information includes recording a video of the display interface and extracting feature information of the video.
[0063] Specifically, the image information of the display interface can be obtained by video recording the display interface. For example, when the status information of the display interface changes, such as when it is detected that a new application is being loaded or switched to a new display interface of the application, the display interface is started to be recorded, and the feature information in the video is extracted. Compared with the form of pictures, videos can obtain more image information, and the image information is continuous, so more comprehensive and continuous feature information can be obtained. For example, the user opens the video application C, and after the video application C is loaded, it enters the homepage of the video application C. The user operates the video application C to enter the personal center of the video application C, opens the playback history, slides to the last played content and clicks to continue playing. According to the embodiment provided by the present invention, when the user opens the video application C, it can be determined that the status information of the display interface has changed. At this time, video recording will be started, the above-mentioned user operations will be recorded, and the feature information in the video will be analyzed. The characteristic information can be the above-mentioned user operation information and operation location information. Therefore, when the user issues an instruction to continue playing, or when it is judged based on big data that the user needs to continue playing, automatic continuous operations will be performed based on the obtained continuous characteristic information. There is no need for the user or the system to issue operation instructions step by step, which improves the user's operating experience and enhances the level of intelligence.
[0064] In some embodiments, when the state information of the display interface stops changing and the duration is greater than or equal to a first preset time, the video recording of the display interface is stopped;
[0065] The extracting feature information of the video includes extracting feature information of key frames in the video.
[0066] Specifically, when it is determined that the status information of the display interface has changed, video recording is started. When the status information change stops and the duration is greater than or equal to the first preset time, it can be considered that the status information change of the display interface has ended. At this time, video recording will be stopped to avoid repeated processing of data and save computing power. Among them, in order to further save computing power, feature information can be extracted only for key frames in the video. The determination of key frames can be system-set, such as extracting a frame in the video at regular intervals and determining it as a key frame. It can also be automatic analysis of big data. When a key part of the picture changes, the video picture at this time can be used as a key frame. By selecting key frames for feature information extraction, processing efficiency can be significantly improved, computing power usage can be reduced, and processing time can be saved.
[0067] In some embodiments, obtaining the status information of the display interface includes obtaining display status information of the application on the display interface;
[0068] If the state information changes, acquiring the image information of the display interface and extracting the characteristic information of the image information includes: if the display state of the application on the display interface changes, and the duration of the changed display state is greater than or equal to a second preset time, starting to acquire the image information of the display interface and extracting the characteristic information of the image information, and when the duration of the changed display state is greater than or equal to a third preset time, stopping acquiring the image information of the display interface;
[0069] Wherein, the second preset time is less than the third preset time;
[0070] The if the display state of the application on the display interface changes includes that the application displayed on the display interface changes or the current display interface of the application changes.
[0071] Specifically, in some embodiments of the present invention, the display interface status information includes information about the application's display status on the display interface, such as full-screen display, half-screen display, or reduced-screen display. Furthermore, to further conserve computing power and avoid unnecessary data processing, in some embodiments of the present invention, when a change in the display interface's display state is detected, the duration of the changed display state is further determined. When the duration is greater than or equal to a second preset time, image information acquisition for the display interface begins. When the duration is greater than or equal to a third preset time, image information acquisition ceases, where the second preset time is less than the third preset time. For example, if a user enters an application and then quickly exits, it could indicate a user error or that user information has been fully acquired and no further processing is required. In this case, image information acquisition for that period may be omitted, thereby conserving computing power and avoiding unnecessary data processing. When the duration of the changed display state is greater than or equal to the second preset time, it can be determined that the user needs to perform a subsequent operation. At this point, image information acquisition and feature extraction from the image information begin. When the duration is greater than or equal to the third preset time, since image acquisition and feature information extraction have continued for a period of time, it can be assumed that the feature information required by the user has been fully extracted. At this point, acquisition of image information for the display interface ceases. For example, when a user enters an application interface for more than 2 seconds, it can be considered that subsequent operations are required and image information acquisition begins. When a user enters an application interface for more than 60 seconds, it can be considered that the feature information has been basically extracted, or the user does not need to perform further operations to control the interface. At this time, image acquisition is stopped.
[0072] In the present invention, the first preset time, the second preset time and the third preset time may be preset by the system or the user, or may be adaptively adjusted and set by the system according to different application scenarios.
[0073] In some embodiments, the feature information includes a text control button area, a graphic control button area, and a text input area;
[0074] The acquiring of the control instruction for the characteristic information includes acquiring a voice instruction, an instruction issued by a server, an instruction transmitted by a third party, or an instruction automatically generated by the system;
[0075] The controlling of the characteristic information according to the control instruction includes performing click, slide, or text input operations on the characteristic information;
[0076] After controlling the characteristic information according to the control instruction, the method further includes feeding back the control result to the user.
[0077] Specifically, the feature information may be operation control area information, such as click, slide, and other operation area information, and may also include text control button areas, graphic control button areas, text input areas, and the like.
[0078] The control instructions may be user voice control instructions, instructions issued by the server, instructions transmitted through a third party (such as control instructions transmitted through the network or a USB flash drive), and control instructions automatically generated by the system.
[0079] Controlling feature information according to control instructions includes operations such as clicking, sliding, and text input.
[0080] The embodiment provided by the present invention further includes providing feedback on the execution result of the control instruction.
[0081] In some embodiments, the feature information further includes coordinate position information on the display interface;
[0082] The method further includes obtaining operation information of a user when controlling the feature information, and simulating the user's operation information when controlling the feature information according to the control instruction;
[0083] The operation information includes click action information, slide action information, and text input action information.
[0084] Specifically, the feature information includes coordinate position information on the display interface, for example, the coordinate position information of the return button on the display interface. Thus, by extracting the coordinate position information, the user's operation can be simulated, for example, simulating the user's sliding, clicking, text input and other operation information.
[0085] The present invention also provides a specific embodiment for a vehicle-mounted central control display screen, such as Figure 2 As shown, steps S201-S204 are included:
[0086] Step S201: The display interface of the vehicle's central control display screen enters a certain system interface.
[0087] Specifically, the control system of the vehicle's central control display screen determines whether the display interface of the current display screen has changed or whether it has entered a certain system interface based on the operation status of the application, and uses this as a trigger condition to trigger subsequent recognition control and other operations.
[0088] Step S202: When the control system of the vehicle-mounted central control display screen detects that the current display interface of the vehicle-mounted central control display screen has entered a certain interface and the stay time exceeds 2 seconds, it can be considered that the user has fully obtained the interface content, and the video recording starts at this time. Otherwise, the recording of the current interface is ignored, and the next interface entered is used as the trigger condition again.
[0089] Specifically, since different users have different perceptions of the interface and the refresh and loading speeds of the interface also vary, the trigger point for video recording will be further accurate to the point where recording will not begin until the content of the current interface is drawn, thereby reducing invalid recording time.
[0090] Step S203: The video is recorded in real time and uploaded to the server synchronously. After the server obtains the locally uploaded video file, it recognizes the image in each second of the video, including determining the interface text or icon content and the position coordinates of the corresponding text or icon in the display interface.
[0091] Specifically, although the overall content of a certain interface may not change much when entering it, sometimes it may be refreshed due to background operations or partial areas of the interface. Therefore, it is necessary to record the entire interface video for all durations and perform text or icon recognition to ensure that even in the above situation, the newly appeared content can be recognized. Among them, the recognition of icons can be converted into text after recognition so as to match the user's voice control instructions. For example, the icon representing return may have multiple display forms, which will be uniformly converted into the text of "return" after recognition. When the user issues a return instruction, it will be matched with the "return" text and controlled.
[0092] Step S204: If the user performs a voice operation, the semantic result is analyzed and matched with the recognition result of the video. If the match is successful, the corresponding semantics are executed and the user is simulated to operate and control the display interface. When the system knows that the current page has been entered for more than 60 seconds and the user has not performed a voice control operation, it means that the user no longer needs to perform further interface control, and the video recording is stopped at this time.
[0093] Specifically, for example, when a user says "play a song," current technology requires API adaptation or debugging of the music application to control the music application through voice control. However, the method provided by an embodiment of the present invention can identify the song play button on the current display interface and its corresponding display interface coordinates, and then complete the voice control by simulating a user click through the system.
[0094] Specifically, such as Figure 3 As shown, the coordinates of the upper left corner and lower right corner of the target recognition area are obtained, and the midpoint position of the area is calculated to obtain the coordinates ((x2+x1) / 2, (y2+y1) / 2). Then, the system simulates a user click to complete the voice operation.
[0095] A second aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the image recognition-based control method of the above embodiment.
[0096] Based on the control method based on image recognition in the above embodiment, the control system based on image recognition proposed in the third aspect of the embodiment of the present invention is described below.
[0097] like Figure 4 As shown, the control system based on image recognition according to the embodiment of the present invention includes:
[0098] The status information acquisition module is used to obtain the status information of the display interface and make judgments.
[0099] The display interface of the present invention can be the display interface of a display screen, or it can be a projected display interface. The status information of the display interface can be the application interface currently displayed on the display interface. For example, the display interface of the present invention is the display interface of a vehicle-mounted central control display screen, and the status information of the display interface is that the current display interface displays an application interface, which can be a map navigation interface, a music broadcast interface, a game interface, etc., or it can be the current interface of an application, such as a music list interface, a music search interface, a music playback interface, etc. of a music player. The status information acquisition module obtains the status information of the display interface and makes a judgment, which can be done by obtaining information about the currently running application, display information of the application on the display interface, information about the current display interface of the application, etc. After obtaining the status information of the display interface, further judgment and analysis are performed.
[0100] The image information acquisition module is used to acquire the image information of the display interface and extract feature information of the image information when the status information changes.
[0101] Specifically, after the status information acquisition module acquires the status information of the display interface, it analyzes whether the status information has changed. For example, after acquiring the application currently being displayed on the display interface or the current display interface information of the application, it is analyzed and compared with the previously acquired status information to determine whether the displayed application has changed or whether the display interface of the application has changed. For example, if the application displayed on the display interface switches from a navigation application to a music playback application, it can be determined that the status information has changed; or if the current display interface of the music playback application changes from a playback interface to a music search interface, it can be determined that the status information has changed. After determining that the status information has changed, the image information acquisition module acquires the image information of the display interface and extracts feature information from the image information. The image information of the display interface can be acquired by taking a screenshot of the current display interface, or by taking a screenshot of only part of the display interface. For example, the display interface may display information of multiple different types and different applications. For example, a vehicle's central control display can simultaneously display weather information, vehicle information (including in-car temperature, battery life, etc.), navigation information, music playback information, and multimedia information (such as WeChat and Weibo information). Some display content, such as vehicle information, is resident, and some applications are not displayed full-screen, occupying only a portion of the display interface. Therefore, when the status information of the display interface changes, such as when a new application is opened but only occupies a portion of the display interface, if the image information of the entire display interface is still obtained and feature information is extracted, redundant information processing will result, resulting in a waste of computing power. At the same time, it may cause unnecessary extension of information processing time, delaying the user's time.
[0102] The feature information of the image information is extracted, and the feature information may include text-based control buttons, graphic-based control buttons, text input areas, and the like. For example, information such as a return icon control button, previous / next page control buttons, fast forward / rewind button information, a progress bar, and a text input box is extracted from the image information. It is understood that the buttons can be circular, rectangular, or other irregularly shaped, or simply in text form. They can also be unconventional buttons, such as the commonly used left-swipe to go to the previous page, right-swipe to go to the next page, left-side up / down swipe to adjust brightness, right-side up / down swipe to adjust volume, and double-click to play / pause. In this case, the buttons of this embodiment can also be unconventional control buttons. Information about unconventional control buttons can be obtained by analyzing information such as the display coordinate area of the application or the application on the display interface, and determining the operation coordinate area of the left-side, double-click, and up-down swipe based on the user's conventional left-right swipe, double-click, and up-down swipe operations, and using the operation coordinate area and the operation as the feature information of the image information. This achieves comprehensive extraction of feature information of the image information on the display interface.
[0103] The control instruction acquisition module is used to acquire the control instruction for the characteristic information.
[0104] In this embodiment of the present invention, the control instruction acquisition module acquires control instructions for feature information. The control instructions may be voice control instructions from the user or control instructions sent from a server, etc. For example, the control instructions for feature information (such as clicks and slides) may be extracted from a voice instruction sent by the user.
[0105] The characteristic information control module is used to control the characteristic information according to the control instruction.
[0106] Specifically, after the control instruction acquisition module obtains the control instruction for the feature information, the feature information control module controls the feature information according to the control instruction. For example, if the user's voice control instruction is to play song B by singer A, then the music search box in the feature information will be searched and played according to the control instruction.
[0107] The image recognition-based control technology solution provided by the present invention can be applied to the control of display interface application software, such as controlling the opening, interface switching, application removal and other operations of music playback software, map navigation software, etc., and can also be applied to the control of display interface content, such as picture processing.
[0108] The system automatically acquires image information from the display interface based on changes in the display interface status information, and then extracts feature information from the image information. After obtaining control instructions for the feature information, the feature information is controlled according to the control instructions. This improves the user experience of image recognition to a certain extent.
[0109] For example, in the prior art, users want to control applications by issuing voice commands. However, application software requires adaptation and other modifications to implement voice control functionality. According to the technical solution provided by the embodiments of the present invention, image recognition technology is used to extract feature information from image information, and then control this feature information based on control instructions. As a result, without the need for adaptive software modifications, it is possible to perform image recognition feature information on most conventional software applications and then control them based on voice control instructions.
[0110] It is understood that some embodiments of the present invention are described using user voice control as an example. However, the application of the present invention is not limited to the field of user voice control technology. It can also be applied to a variety of fields that can replace user operations. For example, when purchasing tickets, it can help users automatically refresh, automatically enter verification codes, automatically place orders and exit, and perform a series of simulated user operations. The image recognition-based control system of the present invention uses image recognition to find feature information and perform control. It does not require customization, adaptive debugging, or modification of the application program. It can control the feature information according to the received control instructions, thereby enhancing the intelligent experience.
[0111] Among them, the control system based on image recognition provided by the present invention obtains the status information of the display interface, and only performs subsequent image information acquisition and feature information extraction processes when the status information changes, thereby avoiding repeated processing of information, reducing the consumption of processor computing power, etc., and also reducing energy consumption.
[0112] The image recognition-based control system provided by an embodiment of the present invention starts to acquire image information and extract feature information when the status information changes. For example, when a new application is starting to load or a display interface of an application is starting to load, it is detected that the status information of the display interface has changed. At this time, the image information is started to be acquired and feature information is extracted, thereby making full use of the period when the application is starting to load or a display interface of an application is starting to load to acquire image information and extract feature information. As a result, image acquisition and feature extraction can be performed in advance on the loaded content.
[0113] In actual operation, extracting feature information from an image requires identifying the image and performing a large number of calculations, which consumes a lot of computing power and also requires a certain amount of computing time. It is not possible to identify and extract immediately. Therefore, the control system based on image recognition provided by an embodiment of the present invention starts to acquire images and extract feature information when it detects that a new application is loading or a new display interface of an application is loading. This can speed up the recognition response speed, greatly save the user's waiting time, and even make the user feel that the time it takes to acquire and extract image information is not needed, thus achieving the purpose of responding to user needs at any time. It avoids interference with the user's normal operations and improves the user experience.
[0114] In some embodiments, the image information acquisition module acquires the image information of the display interface and extracts the feature information of the image information, including that the image information acquisition module acquires the image information of the display interface and compares it with the previously acquired image information, identifies the image area that has changed in the image information, extracts the feature information of the image area and updates it to the previously extracted feature information.
[0115] Specifically, to improve information processing efficiency and avoid redundant information processing, an embodiment of the present invention further provides a solution, including: an image information acquisition module acquiring image information of the display interface and comparing it with previously acquired image information, identifying the image area that has changed in the image information, extracting feature information of the image area, and updating it to the previously extracted feature information. Since in many cases, only part of the displayed content of the display interface has changed, in this case, only the changed part needs to be processed, the feature information of that part needs to be extracted, and it needs to be updated to the previously extracted feature information, thereby greatly reducing the amount of information processing, saving data processing time, reducing user waiting time, and further improving the user experience.
[0116] In some embodiments, the image information acquisition module acquires the image information of the display interface and extracts feature information of the image information, including recording a video of the display interface and extracting feature information of the video.
[0117] Specifically, the image information acquisition module can obtain the image information of the display interface by recording the display interface. For example, when the status information of the display interface changes, such as when it is detected that a new application is loading or switching to a new display interface of the application, the display interface starts to be recorded, and the feature information in the video is extracted. Compared with the form of pictures, videos can obtain more image information, and the image information is continuous, so more comprehensive and continuous feature information can be obtained. For example, the user opens video application C, and after video application C is loaded, it enters the homepage of video application C. The user operates video application C to enter the personal center of video application C, opens the playback history, slides to the last played content and clicks to continue playing. According to the embodiment provided by the present invention, when the user opens the video application C, it can be determined that the status information of the display interface has changed. At this time, video recording will be started, the above-mentioned user operations will be recorded, and the feature information in the video will be analyzed, which may be the above-mentioned user operation information and operation position information. Therefore, when the user issues an instruction to continue playing, or when it is determined based on big data that the user needs to continue playing, automatic continuous operation will be performed based on the obtained continuous feature information, without the need for the user or the system to issue operation instructions step by step, thereby improving the user's operating experience and enhancing the level of intelligence.
[0118] In some embodiments, when the state information of the display interface stops changing and the duration is greater than or equal to a first preset time, the image information acquisition module stops recording the video of the display interface;
[0119] The extracting feature information of the video includes extracting feature information of key frames in the video.
[0120] Specifically, when it is determined that the status information of the display interface has changed, the image information acquisition module starts video recording. When the status information change stops and the duration is greater than or equal to the first preset time, it can be considered that the status information change of the display interface has ended. At this time, the video recording will be stopped to avoid repeated processing of data and save computing power. Among them, in order to further save computing power, feature information can be extracted only for key frames in the video. The determination of key frames can be system-set, such as extracting a frame in the video at regular intervals and determining it as a key frame. It can also be automatic analysis of big data. When a key part of the picture changes, the video picture at this time can be used as a key frame. By selecting key frames for feature information extraction, processing efficiency can be significantly improved, computing power can be reduced, and processing time can be saved.
[0121] In some embodiments, the state information acquisition module acquires the state information of the display interface, including acquiring display state information of the application on the display interface;
[0122] If the state information changes, the state information acquisition module acquires the image information of the display interface and extracts the feature information of the image information, including: if the display state of the application on the display interface changes, and the duration of the changed display state is greater than or equal to a second preset time, then starting to acquire the image information of the display interface and extracting the feature information of the image information; and when the duration of the changed display state is greater than or equal to a third preset time, stopping acquiring the image information of the display interface;
[0123] Wherein, the second preset time is shorter than the third preset time.
[0124] Specifically, in some embodiments of the present invention, the status information of the display interface includes information about the display status of the application on the display interface, such as full-screen display, half-screen display, or reduced-screen display. Furthermore, to further conserve computing power and avoid unnecessary data processing, in some embodiments of the present invention, when the status information acquisition module determines that the display status of the display interface has changed, it further determines the duration of the changed display status. When the duration is greater than or equal to a second preset time, the image information acquisition module begins acquiring image information of the display interface. When the duration is greater than or equal to a third preset time, the image information acquisition module stops acquiring image information, where the second preset time is less than the third preset time. For example, if a user enters an application and then quickly exits, indicating a user error or that user information has been acquired and no further processing is required, image information for that stage may not be acquired, thereby conserving computing power and avoiding unnecessary data processing. When the duration of the changed display status is greater than or equal to the second preset time, it can be determined that the user needs to perform a subsequent operation. At this point, image information acquisition and feature extraction from the image information begin. If the duration is greater than or equal to the third preset time, then, since image acquisition and feature information extraction have continued for a period of time, it can be considered that the feature information required by the user has been extracted, and at this time, acquisition of image information of the display interface will be stopped. For example, if the user has been in an application interface for more than 2 seconds, it can be considered that subsequent operations are required and image information acquisition will begin. If the user has been in an application interface for more than 60 seconds, it can be considered that feature information extraction is basically complete or the user does not need to perform further operations on the interface, and image acquisition will be stopped.
[0125] In the present invention, the first preset time, the second preset time and the third preset time may be preset by the system or the user, or may be adaptively adjusted and set by the system according to different application scenarios.
[0126] In some embodiments, if the display state of the application on the display interface changes, it includes that the application displayed on the display interface changes or the current display interface of the application changes.
[0127] In some embodiments, the feature information includes a text control button area, a graphic control button area, and a text input area;
[0128] The control instruction acquisition module acquires the control instruction for the feature information, including acquiring a voice instruction, an instruction issued by a server, an instruction transmitted by a third party, or an instruction automatically generated by the system;
[0129] The control module controls the characteristic information according to the control instruction, including performing click, slide, and text input operations on the characteristic information;
[0130] After the control module controls the characteristic information according to the control instruction, the system further includes a feedback module for feeding back the control result to the user.
[0131] Specifically, the feature information may be operation control area information, such as click, slide, and other operation area information, and may also include text control button areas, graphic control button areas, text input areas, and the like.
[0132] The control instructions may be user voice control instructions, instructions issued by the server, instructions transmitted through a third party (such as control instructions transmitted through the network or a USB flash drive), and control instructions automatically generated by the system.
[0133] Controlling feature information according to control instructions includes operations such as clicking, sliding, and text input.
[0134] The embodiment provided by the present invention further includes providing feedback on the execution result of the control instruction.
[0135] In some embodiments, the feature information further includes coordinate position information on the display interface;
[0136] The system further includes a user operation information acquisition module for acquiring operation information of the user when controlling the feature information, and the control module simulates the user's operation information when controlling the feature information according to the control instruction;
[0137] The operation information includes click action information, slide action information, and text input action information.
[0138] Specifically, the feature information includes coordinate position information on the display interface, for example, the coordinate position information of the return button on the display interface. Thus, by extracting the coordinate position information, the user's operation can be simulated, for example, simulating the user's sliding, clicking, text input and other operation information.
[0139] The present invention also provides a specific embodiment for a vehicle-mounted central control display screen, such as Figure 5 Shown, including:
[0140] The interface status information acquisition unit acquires interface changes of the display interface and notifies the video recording unit to start and stop video recording according to the completion status of the interface loading.
[0141] Specifically, the interface status information acquisition unit not only monitors changes in the interface, but also provides the user with an interface for setting the response speed of video recording, including the time point to start recording the video and the resolution of the video recording (if the resolution is too high, the data volume is too large and the processing time is too long; if the resolution is too low, the recognition accuracy is not high).
[0142] The video recording unit receives the video recording control signal from the interface status information acquisition unit and communicates with the server at the same time (including uploading the recorded video to the server in real time).
[0143] In addition, the video recording unit is not only responsible for recording the video of the system interface, but also undertakes the management of local video files, regularly deleting the local video cache to avoid excessive occupation of local storage space.
[0144] The video recording unit is also used to receive the image recognition results from the server and store the corresponding results locally for backup.
[0145] When the user performs a voice operation, the voice instruction acquisition unit obtains the user's voice instruction and performs semantic recognition, and the voice instruction execution unit retrieves the image recognition result and executes the user's voice instruction according to the recognition result.
[0146] Specifically, if the current voice command matches the image recognition result, the voice command execution unit executes the corresponding voice control command.
[0147] When executing the user's voice control command, the voice command execution unit simulates the user's operation based on the coordinate information in the image recognition result, such as simulating the user's click, slide, etc.
[0148] Among them, if the current coordinate information has errors and the operation cannot be performed, it will be fed back to the server side for re-labeling training of the video image.
[0149] The feedback unit is responsible for receiving the execution results of the system's voice commands and providing user prompts through text, pictures, voice, ringtones, etc.
[0150] Specifically, feedback prompts are provided based on different scenarios. For example, if the user is currently playing music with background music, and the current voice command would cause auditory or visual changes, the voice feedback reminder will be primarily text, without sound. Similarly, if the current voice command would cause a physical change (such as temperature or air volume), the text reminder will be the primary notification, allowing the user to perceive the change in voice control without sound. This avoids interruptions to the user and improves the user experience.
[0151] The image recognition-based control system provided by the embodiment of the present invention obtains image data in the video stream according to video recording, and automatically recognizes text and the coordinates of the corresponding text without the need for user clicks. After the user voice command is issued, the command is identified as directly identical or related to the current interface content, and the corresponding voice operation is performed. For example, after opening a video application, if the interface contains the text "TV series", the user can directly issue the voice command "open TV series" to open it, simulating the user's click operation and facilitating user operation. It is particularly practical when the user is driving, and can be adapted to various applications without the need to reopen for adaptation.
[0152] A fourth aspect of the present invention provides a vehicle, such as Figure 6 As shown, the vehicle of the embodiment of the present invention includes a display device and a control system based on image recognition in the embodiment. For example, the display device may include a vehicle-mounted central control display screen, a head-up display HUD, etc.
[0153] According to the vehicle of the embodiment of the present invention, by adopting the control system based on image recognition of the above embodiment, there is no need to perform adaptive debugging on the software, and the user can also perform voice control on the software, thereby improving the user experience.
[0154] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "example," "specific example," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.
[0155] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A control method based on image recognition, characterized in that: include: Obtain the status information of the display interface and make judgments; If the state information changes, image information of the display interface is obtained and feature information of the image information is extracted; obtaining the image information of the display interface and extracting the feature information of the image information includes recording a video of the display interface and extracting feature information of the video, and when the state information of the display interface stops changing and the duration is greater than or equal to a first preset time, stopping the video recording of the display interface, the feature information includes operation information and operation position information, wherein the change in state information includes: a change in the application displayed on the display interface or a change in the current display interface of the application, and the operation information includes click action information, sliding action information, and text input action information; extracting the feature information of the video includes extracting feature information of key frames in the video; Obtaining a control instruction for the characteristic information; Controlling the characteristic information according to the control instruction includes: automatically and continuously operating the acquired continuous characteristic information according to the control instruction without issuing operation instructions step by step; Wherein, the acquiring of the status information of the display interface includes acquiring the display status information of the application on the display interface; If the state information changes, acquiring the image information of the display interface and extracting the characteristic information of the image information includes: if the display state of the application on the display interface changes, and the duration of the changed display state is greater than or equal to a second preset time, starting to acquire the image information of the display interface and extracting the characteristic information of the image information, and when the duration of the changed display state is greater than or equal to a third preset time, stopping acquiring the image information of the display interface; Wherein, the second preset time is shorter than the third preset time.
2. The control method based on image recognition according to claim 1, characterized in that: The characteristic information includes a text control button area, a graphic control button area, and a text input area; The acquiring of the control instruction for the characteristic information includes acquiring a voice instruction, an instruction issued by a server, an instruction transmitted by a third party, or an instruction automatically generated by the system; The controlling of the characteristic information according to the control instruction includes performing click, slide, or text input operations on the characteristic information; After controlling the characteristic information according to the control instruction, the method further includes feeding back the control result to the user.
3. The control method based on image recognition according to claim 1, characterized in that: The characteristic information includes coordinate position information on the display interface; The method further includes acquiring operation information of a user when controlling the feature information, and simulating the user's operation information when controlling the feature information according to the control instruction.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the image recognition-based control method according to any one of claims 1 to 3 is implemented.
5. A control system based on image recognition, characterized in that: include: Status information acquisition module, used to obtain status information of the display interface and make judgments; An image information acquisition module, configured to acquire image information of the display interface and extract feature information of the image information when the status information changes; The image information acquisition module acquires the image information of the display interface and extracts feature information of the image information, including: recording the display interface and extracting feature information of the video; when the state information of the display interface stops changing and the duration is greater than or equal to a first preset time, the image information acquisition module stops recording the video of the display interface, the feature information includes operation information and operation position information, wherein the change in the state information includes: a change in the application displayed on the display interface or a change in the current display interface of the application; the operation information includes click action information, sliding action information, and text input action information; and extracting the feature information of the video includes extracting feature information of key frames in the video; A control instruction acquisition module, used to acquire control instructions for the feature information; A characteristic information control module, configured to control the characteristic information according to the control instruction, including: automatically and continuously operating the acquired characteristic information according to the control instruction without issuing operation instructions step by step; The state information acquisition module acquires the state information of the display interface, including acquiring display state information of the application on the display interface; If the state information changes, the state information acquisition module acquires the image information of the display interface and extracts the feature information of the image information, including: if the display state of the application on the display interface changes, and the duration of the changed display state is greater than or equal to a second preset time, then starting to acquire the image information of the display interface and extracting the feature information of the image information; and when the duration of the changed display state is greater than or equal to a third preset time, stopping acquiring the image information of the display interface; Wherein, the second preset time is shorter than the third preset time.
6. The control system based on image recognition according to claim 5, characterized in that: The characteristic information includes a text control button area, a graphic control button area, and a text input area; The control instruction acquisition module acquires the control instruction for the feature information, including acquiring a voice instruction, an instruction issued by a server, an instruction transmitted by a third party, or an instruction automatically generated by the system; The control module controls the characteristic information according to the control instruction, including performing click, slide, and text input operations on the characteristic information; After the control module controls the characteristic information according to the control instruction, the system further includes a feedback module for feeding back the control result to the user.
7. The control system based on image recognition according to claim 5, characterized in that: The characteristic information includes coordinate position information on the display interface; The system further includes a user operation information acquisition module for acquiring the user's operation information when controlling the feature information, and the control module simulates the user's operation information when controlling the feature information according to the control instruction.
8. A vehicle, characterized in that: The system comprises a display device and a control system based on image recognition as described in any one of claims 5 to 7.
Citation Information
Patent Citations
Touch screen operation method and touch screen terminal
CN106980457A
Voice control method, terminal equipment, cloud server and system
CN108538291A
Speech central control method and device based on image recognition
CN109471678A