Anaphora resolution method based on spatial position relation, computer equipment and storage medium
By constructing a spatial location relationship database and performing semantic search and vectorization processing, the problem of insufficient precise vocabulary and spatial understanding ability of in-vehicle voice assistants has been solved, improving interaction efficiency and accuracy, and providing an intelligent and convenient user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 南昌勤胜电子科技有限公司
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-10
AI Technical Summary
Existing in-vehicle voice assistants rely on precise vocabulary and lack spatial understanding capabilities, making it difficult to understand natural language descriptions and vague references based on relative location, resulting in low interaction efficiency.
By extracting spatial location relationship features and visual description features from the query information, a spatial location relationship database is constructed, and semantic search and vectorization processing are performed to improve the system's spatial understanding ability and referential resolution ability.
It achieves accurate recognition and understanding of users' natural language descriptions, improves the adaptability and interaction efficiency of in-vehicle voice assistants in complex query scenarios, and provides a more intelligent, convenient and natural human-computer interaction experience.
Smart Images

Figure CN121833732A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, and in particular to a reference resolution method based on spatial position relationship, a computer device and a storage medium. BACKGROUND
[0002] With the rapid development of intelligent cockpit technology, the vehicle-mounted voice assistant has become a mandatory artificial interaction function of modern cars, greatly improving the user's driving experience and operation convenience. Users can query the vehicle-mounted control function and operate various functions of the vehicle through the vehicle-mounted voice assistant.
[0003] However, the current vehicle-mounted voice assistant still has many limitations in practical application, mainly in the following aspects:
[0004] (1) Dependence on accurate vocabulary: users must know the accurate name of the control (such as "double flash button", "air conditioner knob") to effectively query. When the user faces a control that he does not know, the query function cannot function properly, causing great inconvenience to the user.
[0005] (2) Lack of spatial understanding ability: In daily communication, people often use natural language descriptions based on relative positions to refer to a certain object, such as "the one on the left of the steering wheel" and "the button below the center control screen". This kind of reference based on spatial position relationship is a common phenomenon in human communication and conforms to people's language habits. However, the current vehicle-mounted voice assistant is difficult to directly understand this kind of natural language description.
[0006] (3) Weak reference resolution ability: when using a vague description to inquire (such as "what is that round button"), it is impossible to determine the specific control object referred to by the user in combination with the context and spatial position relationship.
[0007] Therefore, it is urgent to improve the existing technology to improve the interaction efficiency and user experience and meet the increasingly diverse needs of users.
[0008] The above information is given as background information only to assist in understanding the present application, and does not determine or acknowledge whether any of the above content can be used as prior art against the present application. SUMMARY
[0009] The present application provides a reference resolution method based on spatial position relationship, a computer device and a storage medium to improve the interaction efficiency and user experience.
[0010] To achieve the above purpose, the present application provides the following technical solutions:
[0011] In a first aspect, the present application provides a method for resolving reference based on spatial position relationship, which comprises:
[0012] receiving user inputted query information to be resolved;
[0013] extracting keywords from the query information, the keywords including spatial position relationship features and visual description features;
[0014] combining the extracted keywords into a query sentence;
[0015] performing semantic search in a preset spatial position relationship database according to the query sentence to obtain search results;
[0016] responding to the user according to the search results.
[0017] Further, in the method for resolving reference based on spatial position relationship, after the step of combining the extracted keywords into a query sentence, the method further comprises:
[0018] vectorizing the query sentence to obtain a query sentence vector;
[0019] Correspondingly, the step of performing semantic search in a preset spatial position relationship database according to the query sentence to obtain search results is:
[0020] performing semantic vector search in a preset spatial position relationship vector database according to the query sentence vector to obtain search results.
[0021] Further, in the method for resolving reference based on spatial position relationship, the step of responding to the user according to the search results comprises:
[0022] determining whether the search results are unique;
[0023] if yes, generating reference resolution data according to the search results to respond to the user;
[0024] if no, providing the search results one by one to the user for confirmation, and after obtaining the user confirmation, generating reference resolution data according to the confirmed search results to respond to the user.
[0025] Further, in the method for resolving reference based on spatial position relationship, the step of providing the search results one by one to the user for confirmation, and after obtaining the user confirmation, generating reference resolution data according to the confirmed search results to respond to the user comprises:
[0026] providing the search results one by one to the user for confirmation in a multi-modal manner;
[0027] After obtaining the user confirmation, according to the confirmed search result, the resolution data is generated to respond to the user.
[0028] Further, in the method for resolving the reference based on the spatial position relationship, the method further comprises:
[0029] A spatial position relationship database is constructed, and the spatial position relationship database stores the name, visual description, spatial position relationship and function information of each spatial object.
[0030] Further, in the method for resolving the reference based on the spatial position relationship, the method further comprises:
[0031] A spatial position relationship database is constructed, and the spatial position relationship database stores the name, visual description, spatial position relationship and function information of each spatial object.
[0032] The text information in the spatial position relationship database is vectorized to obtain a text vector.
[0033] The text vector is stored in a database to construct a spatial position relationship vector database.
[0034] Further, in the method for resolving the reference based on the spatial position relationship, the spatial position relationship feature comprises at least one of an absolute position relationship and a relative position relationship.
[0035] Further, in the method for resolving the reference based on the spatial position relationship, the visual description feature comprises at least one of an icon, shape, color and type.
[0036] In a second aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method for resolving the reference based on the spatial position relationship according to the first aspect when executing the computer program.
[0037] In a third aspect, the present application provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are executed by a computer processor to implement the method for resolving the reference based on the spatial position relationship according to the first aspect.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] The application provides a referential resolution method based on spatial position relations, a computer device and a storage medium, which extracts spatial position relation features and visual description features in query information, and performs semantic search in combination with a preset spatial position relation database, effectively solves the problems that an existing vehicle-mounted voice assistant relies on accurate vocabulary, lacks spatial understanding ability and has weak referential resolution ability when performing human-computer interaction, realizes accurate recognition and understanding of natural language description of a user, significantly improves adaptability, accuracy and interaction efficiency of the vehicle-mounted voice assistant in a complex query scene, provides a more intelligent, convenient and natural interaction experience for the user, thereby greatly improving practicality and user satisfaction of the vehicle-mounted voice assistant, and has obvious practical value and wide application prospect.
[0040] The present application has other characteristics and advantages, which will be apparent from and / or set forth in the accompanying drawings and the following detailed description, which together serve to explain certain principles of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0042] Figure 1 is one of the flowcharts of the referential resolution method based on spatial position relations provided by the first embodiment of the present application;
[0043] Figure 2 is one of the flowcharts of the referential resolution method based on spatial position relations provided by the first embodiment of the present application;
[0044] Figure 3 is one of the flowcharts of the further refined S105 provided by the first embodiment of the present application;
[0045] Figure 4 is the second flowchart of the further refined S105 provided by the first embodiment of the present application;
[0046] Figure 5 is the flowchart of S201 provided by the first embodiment of the present application;
[0047] Figure 6 is the flowchart of S301-S303 provided by the first embodiment of the present application;
[0048] Figure 7 is a schematic diagram of a spatial position relationship database provided by an embodiment of the present application;
[0049] Figure 8 is a structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0050] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0051] Embodiment one
[0052] Please refer to Figure 1 is a flow schematic diagram of a reference resolution method based on a spatial position relationship provided by an embodiment of the present application. The method is suitable for scenarios such as a user using a vehicle-mounted voice assistant to query the function of a vehicle-mounted control, and the method can be implemented by software and / or hardware. The method specifically includes the following steps:
[0053] S101, receiving query information input by a user and to be resolved in reference;
[0054] It should be noted that this is the starting step of the entire process. In the intelligent cabin, the system can receive information input by the user through the voice input by the vehicle-mounted voice assistant, such as information about the function of the vehicle-mounted control that the user wants to query. These information are the basic data for subsequent reference resolution processing.
[0055] For example, the query information is "what is the function of the knob on my left side that draws a fan", which is a typical query information to be resolved in reference, which contains the user's vague description of a certain vehicle-mounted control, but does not explicitly mention the accurate name of the control.
[0056] S102, extracting keywords from the query information, the keywords including a spatial position relationship feature and a visual description feature;
[0057] It should be noted that this step is to deeply analyze the query information input by the user and to filter out the words with key significance. These words are mainly divided into two categories, namely the spatial position relationship feature and the visual description feature. These two features are crucial for accurately understanding the vehicle-mounted control referred to by the user.
[0058] The spatial position relationship feature includes at least one of an absolute position relationship and a relative position relationship. The absolute position relationship refers to a specific position of an object in a specific coordinate system, such as "the second one on the right side of the steering wheel". The relative position relationship refers to a position relationship between objects, such as "next to the volume adjustment knob".
[0059] The visual description feature includes at least one of an icon, a shape, a color, and a type. The icon refers to a graphical symbol on a control for identifying its function, such as a speaker icon representing a speaker function and a phone icon representing a phone function. The shape refers to the geometric shape of the control, such as a circular button and a square knob. The color refers to the appearance color of the control, such as a red button and a blue knob. The type refers to the category to which the control belongs, such as a button, a knob, and a dial.
[0060] For example, in the query information "what is the function of the knob with a fan drawn on the left side of my hand", the extracted spatial position relationship features include "left side of the hand" and "on top", and the extracted visual description features include "fan" and "knob".
[0061] S103, combining the extracted keywords into a query statement;
[0062] It should be noted that this step combines the spatial position relationship features and the visual description features extracted in S102 into a complete query statement according to certain logic and order. This query statement will be used as the basis for subsequent searching in the database.
[0063] For example, the spatial position relationship features "left side of the hand" and "on top" and the visual description features "fan" and "knob" are combined into a query statement "left side of the hand fan knob". Through this combination, the features of the vehicle-mounted control described by the user can be more accurately and comprehensively expressed.
[0064] S104, performing semantic search in a pre-set spatial position relationship database according to the query statement to obtain a search result;
[0065] It should be noted that this step uses the query statement generated in S103 to search in a pre-constructed spatial position relationship database. This database stores a large amount of spatial position relationship and related visual description features of vehicle-mounted controls. Through semantic search, the system can find records matching the query statement from the database, and these records are the search results.
[0066] For example, assume that the spatial position relationship database stores detailed information of various in-vehicle controls, including their positions in the vehicle, appearance features, etc. When the query statement "left-hand side fan knob" is input, the system searches the database for in-vehicle control information that matches "position on the left-hand side, has a fan icon, and is a knob type", and returns the found information as the search result.
[0067] S105, according to the search result, answering the user.
[0068] It should be noted that this step is an analysis and processing of the search result obtained in S104, and the information related to the user query is fed back to the user in a suitable manner, completing the entire query interaction process.
[0069] For example, if the search result finds the in-vehicle control described by the user, the system will inform the user of the function of the control, such as "the left-hand side knob with a fan picture you described is the air conditioner air volume knob, which is used to rotate to adjust the air volume of the air conditioner".
[0070] The embodiment of the present application has significant beneficial effects, which can effectively solve the limitations of the existing in-vehicle voice assistant, and improve user experience and interaction efficiency.
[0071] Firstly, the method extracts the spatial position relationship features and visual description features in the query information, so that the user can effectively query through descriptive language without knowing the exact name of the control when querying the function of the in-vehicle control. For example, the user can use natural language descriptions such as "the one on the left side of the steering wheel" and "the button below the center screen" to refer to the control, and the system can accurately understand and recognize these descriptions, thereby solving the problem of limited query function caused by dependence on precise vocabulary in the prior art.
[0072] Secondly, the method enhances the spatial understanding ability of the in-vehicle voice assistant. In daily communication, people are used to using relative position-based description to refer to objects, and the present application can directly understand such natural language descriptions and match the user's description with specific control objects, thereby realizing more natural and intuitive human-computer interaction.
[0073] In addition, the method also significantly improves the reference resolution capability. When the user uses a vague description to inquire, the system can determine the specific control object referred to by the user by combining the context and spatial position relationship. For example, for a vague description such as "what is that round button", the system can accurately identify and answer the user's question through semantic search and matching of the spatial position relationship database, thereby improving the adaptability, accuracy and interaction efficiency of the in-vehicle voice assistant in complex query scenarios.
[0074] To sum up, the application comprehensively improves the intelligent level of the vehicle-mounted voice assistant through innovative technical means, so that it can better meet the diversified needs of users in actual driving scenarios, provide users with more intelligent, convenient and natural interaction experience, and has significant practical value and broad application prospect.
[0075] Please refer to Figure 2 In an embodiment of the present embodiment, after S103, the method further adds a key step, namely S103.5, and the S104 is adaptively adjusted, which aims to further improve the performance and search effect of the method.
[0076] S103.5, vectorizing the query statement to obtain a query statement vector;
[0077] It should be noted that the core of this step is to convert the combined query statement from natural language text form to vector form. Vector is a mathematical representation, in the computer field, through specific algorithms and models, the semantic, syntax and other information in the text can be encoded into vectors. Vectorization processing is of great significance, on the one hand, the speed of vector in computer calculation and processing is much faster than that of text form, which can significantly improve the efficiency of the whole process; on the other hand, vectorization can more accurately capture the semantic information in the query statement, thereby improving the accuracy of subsequent search matching, so that the system can more accurately understand the query intention of the user.
[0078] For example, assuming that the combined query statement is "left hand fan knob", through the vectorization algorithm, the statement is converted into a multi-dimensional vector, and each dimension of the vector contains different words in the statement and their combined semantic features.
[0079] Correspondingly, the S104 adaptively adjusts to the following steps:
[0080] S104, according to the query statement vector, performing semantic vector search in a preset spatial position relationship vector database to obtain a search result.
[0081] It should be noted that the adjusted S104 step no longer uses the original query statement for search, but uses the query statement vector obtained by the S103.5 step to search in the pre-constructed spatial position relationship vector database. This database stores vehicle control related information processed by vectorization, including their spatial position relationship, visual description features, etc. are also presented in the form of vectors. Semantic vector search finds the most matching record with the query intention by calculating the similarity between the query statement vector and the vector in the database, so as to obtain the search result.
[0082] Comparison with traditional search methods:
[0083] Traditional keyword search: Traditional keyword search mainly relies on exact matching, that is, the keywords in the query statement must be completely consistent with the records in the database to match successfully. Even if it is semantic search, it only depends on the degree of semantic expansion, that is, the number of expanded synonyms. If the range of synonym expansion is limited, the accuracy of the search will still be greatly limited. For example, when the user queries "the button in the circle", if the database only has the record of "the round button" and there is not enough synonym expansion, the traditional search method may not be able to associate the two, resulting in no relevant results found.
[0084] Advantages of semantic vector search: Semantic vector search converts text into vector representation and calculates similarity, realizing deep understanding of query intent. It can not only capture the semantic association between words, but also return relevant results even if the query statement and the document in the database use different words. Moreover, in the vector space, through similarity calculation, it can tolerate a certain degree of deviation, map the near-synonymous query to the similar vector area, and thus return accurate results. For example, for the wrong expression "phone icon", semantic vector search can recognize its semantic similarity with "call icon", correct it and match it to the correct result, greatly improving the accuracy and flexibility of the search.
[0085] In summary, by adding the S103.5 step to vectorize the query statement and adjusting the S104 step for semantic vector search, the reference resolution method based on spatial position relationship in this embodiment has been significantly improved in terms of computing processing speed, search matching accuracy, and understanding ability of query intent. It can better meet the diverse query needs of users, especially when dealing with ambiguous, synonymous or incorrect expressions, it can still accurately return relevant results, providing users with a better in-vehicle voice interaction experience, and further enhancing the practicality and competitiveness of the method.
[0086] Please refer to Figure 3 In an embodiment of the present embodiment, the S105 can be further refined to include the following sub-steps, different processing strategies are adopted for different search results to finally generate accurate reference resolution data and complete the response to the user, ensuring the completeness and effectiveness of the entire reference resolution process:
[0087] S1051, determine whether the search result is unique; if yes, execute S1052, if no, execute S1053;
[0088] It should be noted that this step is a preliminary screening and judgment of the search results obtained by semantic vector search. Due to the actual situation, there may be multiple cases leading to non-unique search results. For example, there may be multiple controls in the vehicle system that have similar appearance, similar function, and easily confused spatial position description; or the semantics of the user query statement has certain ambiguity (such as the spatial position relationship is not specific enough), so that multiple vehicle controls meet the search conditions. By judging whether the search results are unique, the subsequent processing flow can be divided into two different paths to improve the pertinence and efficiency of processing.
[0089] For example, assuming that the user queries "the button on the right side that can adjust the temperature", after semantic vector search, two search results may be obtained, one is a button on the right side for adjusting the temperature of the driver's seat, and the other is a button on the right side for adjusting the temperature of the front passenger's seat. At this time, the search results are not unique; if only the button on the right side for adjusting the temperature of the driver's seat meets the query condition, then the search results are unique.
[0090] S1052、According to the search results, generate reference resolution data and respond to the user;
[0091] It should be noted that when the search results are unique, it means that the system has accurately found the vehicle control that matches the user's query intent. At this time, the system can directly generate corresponding reference resolution data according to this unique search result. The reference resolution data usually contains detailed information of the specific vehicle control referred to by the user query, such as name, function, position, etc. Then, the system organizes these information in a suitable way to respond to the user and inform the user of the specific object referred to by the user query.
[0092] For example, continuing the above example, if the search results are unique, that is, it is determined that it is the button on the right side for adjusting the temperature of the driver's seat, the reference resolution data generated by the system may include "You query the driver's seat temperature adjustment button on the right side. The button can adjust the temperature of the driver's seat area." Then the system transmits this response information to the user through voice or text, etc.
[0093] S1053, provide the search results one by one to the user for confirmation, and after obtaining the user's confirmation, generate reference resolution data according to the confirmed search results and respond to the user.
[0094] It should be noted that when it is judged that the search result is not unique, it means that there are multiple vehicle controls that may meet the user's query intention. In order to avoid system automatic selection of possible errors and improve the accuracy of anaphora resolution, the system will present these search results one by one to the user for confirmation. The user can select the result that best meets his query intention from multiple options according to his actual needs and intentions. After obtaining the user's confirmation, the system generates anaphora resolution data according to the user's confirmed search result, and responds to the user in a similar manner to S1052.
[0095] For example, as mentioned earlier, the user query "the button on the right side that can adjust the temperature" gets two search results, the system will present these two results to the user respectively, for example, with a voice prompt "we found two buttons that may meet your needs, the first is the button to adjust the temperature of the driver's seat, located on the right side; the second is the button to adjust the temperature of the front passenger seat, also located on the right side, please confirm which one you want". After the user confirms through voice reply or selection operation, the system generates anaphora resolution data and responds according to the user's confirmed result, for example, if the user confirms the button to adjust the temperature of the driver's seat, the system responds "you confirm the driver's seat temperature adjustment button located on the right side, which can adjust the temperature of the driver's seat area".
[0096] In summary, by refining S105, the anaphora resolution method based on spatial position relationship in this embodiment can more flexibly and accurately handle search results in different situations. Whether the search result is unique or not, the finally generated anaphora resolution data can be ensured to be accurate, thereby providing accurate responses for the user, effectively solving the problem of ambiguous reference in the user query process, and improving the practicality and user experience of the vehicle voice interaction system.
[0097] Please refer to Figure 4 , in an embodiment of this embodiment, S1053 is further refined on the basis of Figure 3 . In Figure 3 , when the search result is not unique, only the search results are simply mentioned to be provided one by one to the user for confirmation, but the way is not explicitly provided. The Figure 4 improvement is that the user confirms the search result by introducing a multi-modal way, which can greatly improve the accuracy and experience of user confirmation, and ensure that the entire anaphora resolution process is more perfect and reliable. The following is a detailed description of each sub-step after refining S1053.
[0098] S10531, provide the search results one by one to the user for confirmation through a multi-modal way;
[0099] It should be noted that the multi-modal mode refers to comprehensively using text, image, voice and other forms to present the search results. Different modalities of information have different characteristics and advantages, and can convey detailed information of the search results to the user from multiple dimensions, helping the user to more comprehensively and accurately understand the characteristics of the vehicle control represented by each search result, so as to make correct confirmation selection.
[0100] In practical application, the three modalities can complement each other. For example, first introduce the search results through voice broadcast, and at the same time display the corresponding text information and image on the screen. The user can listen to the voice while looking at the text and image, and comprehensively understand the search results from different angles, greatly improving the accuracy and efficiency of confirmation. For example, the user queries "the control switch for lighting in the car", and the system presents the search results in a multi-modal way, and the voice broadcast is "the first search result is a long strip-shaped button with a light icon on the left side of the driver's seat, which is used to control the reading light in the car". At the same time, the text description and actual picture of the button are displayed on the screen, so that the user can more accurately determine whether it is the control that he wants to confirm.
[0101] S10532、In response to the user's confirmation, the system generates resolution data according to the confirmed search result, and answers the user.
[0102] It should be noted that when the user confirms the search results through the multi-modal mode and makes a selection, the system generates corresponding resolution data according to the search result confirmed by the user. The resolution data contains detailed information of the specific vehicle control referred to by the user's query, which is an accurate response to the user's query intention. Then, the system organizes these resolution data in a suitable way to answer the user and inform the user of the specific object referred to by the user's query and related information.
[0103] For example, continuing the above example, the user confirms that the switch located on the left side of the driver's seat for controlling the reading light in the car is the query target. The resolution data generated by the system may include "You query the reading light control switch located on the left side of the driver's seat, which is a long strip-shaped button with a light icon, and pressing it can turn on or off the reading light in the car". The system can convey this answer information to the user through voice or text, and complete the entire resolution process.
[0104] In summary, by further refining S1053 to adopt a multi-modal manner to let the user confirm the search result, and generating a reference resolution data response to the user based on the confirmation result, the reference resolution method based on spatial position relationship in this embodiment is more scientific, accurate and humanized when dealing with the case where the search result is not unique. The multi-modal manner makes full use of the advantages of different modal information, helps the user to confirm the search result from multiple angles, effectively avoids the problem of insufficient information or misunderstanding caused by a single modal, greatly improves the accuracy of reference resolution and user satisfaction, and further improves the performance and practicality of the vehicle-mounted voice interaction system.
[0105] In different embodiments of this embodiment, according to the different of the base graph (S101) Figure 1 or Figure 2 ), the reference resolution method based on spatial position relationship is further expanded. On the basis of Figure 1 , by adding the step of constructing a spatial position relationship database (i.e., S201), a richer data basis is provided for subsequent possible reference resolution operations; on the basis of Figure 2 , not only the spatial position relationship database (i.e., S301) is constructed, but also the text information in the database is further vectorized (i.e., S302) and a spatial position relationship vector database (i.e., S303) is constructed, so that the data can better adapt to subsequent processing methods such as vector search, and the efficiency and accuracy of reference resolution are improved. The following is a detailed description of each newly added step.
[0106] Please refer to Figure 5 , on the basis of Figure 1 , the method can further include the following steps:
[0107] S201, constructing a spatial position relationship database, the spatial position relationship database stores the name, visual description, spatial position relationship and function information of each spatial object.
[0108] It should be noted that constructing a spatial position relationship database is the basic preparation work of the entire reference resolution method. This database aims to comprehensively record the multi-dimensional information of each spatial object, so that relevant information can be quickly and accurately obtained for matching and reasoning when processing user queries subsequently. The name is used to identify the spatial object; the visual description helps the user to identify the object from the appearance; the spatial position relationship clearly shows the relative position of the object in space, which is crucial for reference resolution based on spatial position; and the function information explains the purpose of the object, which helps to answer the function of the user query.
[0109] For example, in the constructed in-vehicle space position relationship database, a space object can be a "wiper control lever". Its name is "wiper control lever"; the visual description can be "located on the right side of the steering wheel, in the form of a long rod, with different gear marks"; the spatial position relationship can be "left side is the steering wheel, right side is the light control lever, upper side is the instrument panel, and lower side is empty"; and the function information can be "by up and down lever adjustment control, different working gears of the wiper can be adjusted, such as intermittent gear, low speed gear, high speed gear, etc., for removing rain or debris on the windshield".
[0110] Please refer to Figure 6 , on the basis of Figure 2 , the method can further include the following steps:
[0111] S301, constructing a space position relationship database, the space position relationship database stores the name, visual description, spatial position relationship and function information of each space object;
[0112] It should be noted that, similar to S201, this step is also to construct a comprehensive space position relationship database to provide data support for subsequent vector processing and anaphora resolution. In the in-vehicle scenario, various in-vehicle controls are space objects, and the detailed information of the space objects is helpful to accurately understand the objects referred to by the user query.
[0113] For example, as shown in Figure 7 , taking the in-vehicle control "hazard warning light button" as an example, its name is recorded in the database as "hazard warning light button"; the visual description is "red triangular icon button"; the spatial position relationship is "the upper space object is the air outlet, the lower space object is the central control screen, the left space object is empty, and the right space object is the air volume knob"; and the function information is "pressing the button turns on the double flash light, which serves as an emergency warning function". Through such detailed recording, the various characteristics of the space object can be clearly described.
[0114] S302, vectorizing the text information in the space position relationship database to obtain a text vector;
[0115] It should be noted that through vectorization, the semantic information of the text can be encoded into a vector form, so that texts with similar semantics are close in the vector space, which is convenient for subsequent vector search and similarity calculation operations. In this embodiment, the name, visual description, spatial position relationship and function information in the space position relationship database are vectorized, which can convert these text information into a vector form that can be used for calculation and comparison.
[0116] For example, taking the function information of the "hazard warning light button" as an example, the text is converted into a multi-dimensional numerical vector through a specific text vectorization algorithm. Each dimension in the vector represents a certain semantic feature in the text. By calculating the distance or similarity between vectors, the semantic similarity between different function information can be determined.
[0117] S303, store the text vector into a database to construct a spatial position relationship vector database.
[0118] It should be noted that the text vector obtained through vectorization processing is stored in a special database to form a spatial position relationship vector database. This database is related to the previously constructed spatial position relationship database, but stores the data in the form of vectorization. The construction of the spatial position relationship vector database provides a data basis for the subsequent reference resolution method based on vector search, so that the system can quickly and accurately find the most matching object from a large number of spatial objects based on the user's query semantics.
[0119] For example, the name vector, visual description vector, spatial position relationship vector and function information vector of the "hazard warning light button" are stored in the spatial position relationship vector database. When the user queries, the system can also vectorize the user's query statement, and then search in the spatial position relationship vector database to find the most similar spatial object vector to the query vector, so as to determine the vehicle control referred to by the user query.
[0120] In summary, by expanding the method based on different basic graphs, constructing the spatial position relationship database and further constructing the spatial position relationship vector database, the reference resolution method based on spatial position relationship of the embodiment can more comprehensively and accurately process user queries. The spatial position relationship database provides rich spatial object information, and the spatial position relationship vector database enables this information to be stored and searched in a form that is easy for computers to process, greatly improving the efficiency and accuracy of reference resolution.
[0121] Although the terms such as query information and keywords are used more in the present application, the possibility of using other terms is not excluded. The use of these terms is only to facilitate the description and explanation of the essence of the present application; any additional limitation is contrary to the spirit of the present application.
[0122] Embodiment two
[0123] Figure 8 A structural schematic diagram of a computer device provided for embodiment two of the present application. Figure 8A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present invention is shown. Figure 8 The computer device 12 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0124] like Figure 8 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0125] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0126] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0127] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 8 Not shown; usually referred to as a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0128] Program / utility 40 having a set of program modules 42 can be stored in memory 28 by way of example, such program modules 42 include an operating system, one or more application programs, other program modules, and program data, each or some combination thereof, which can include implementations of the network environment as in each of the above examples or some combination thereof. Program modules 42 generally carry out the functions and / or methodologies of embodiments of the application as described herein.
[0129] Computer device 12 can also communicate with one or more external devices 14 such as a keyboard or pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer device 12; and / or any devices (e.g., network card, modem, etc.) that enable computer device 12 to communicate with one or more other computing devices. Such communication can occur via input / output (I / O) interface(s) 22. Still yet, computer device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or the Internet) through network adapter 20. As an example, network adapter 20 can include a modem, a network card (wireless or wired), or other well-known interface devices. It is noted that these devices can be external to computer device 12, such as in the case of a network card, or can be internal to computer device 12, such as in the case of a modem or other device. Figure 8 It is to be appreciated that other hardware and / or software modules that can be used in conjunction with computer device 12 can also be utilized with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0130] Processing unit 16 performs various function applications and data processing by running programs stored in system memory 28, such as implementing the spatial position relationship based anaphora resolution method provided by embodiments of the present application.
[0131] Embodiment Three
[0132] Embodiment three of the present application provides a computer readable storage medium, which stores computer executable instructions, the instructions being executed by a processor to implement the spatial position relationship based anaphora resolution method provided by all embodiments of the present application.
[0133] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0134] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0135] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0136] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In an embodiment, the present application is directed to computer program products comprising machine-readable media for carrying or having machine-executable instructions or programs
[0137] Finally, it should be noted that the above embodiments have been described in the specification and drawings of the present application, but this does not limit the patent protection scope of the present application. Any equivalent structure or equivalent process replacement or modification based on the essential concept of the present application, using the content described in the specification and drawings of the present application, and directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, etc., are all included in the patent protection scope of the present application.
Claims
1. A method for resolving anaphora based on spatial location relations, characterized by, The method comprises: receiving user input query information to be anaphora resolution; extracting keywords from the query information, the keywords including spatial location relationship features and visual description features; combining the extracted keywords into a query statement; performing semantic search in a preset spatial location relationship database according to the query statement to obtain a search result; responding to the user according to the search result.
2. The method of claim 1, wherein, After the step of combining the extracted keywords into a query statement, the method further comprises: vectorizing the query statement to obtain a query statement vector; Correspondingly, the step of performing semantic search in a preset spatial location relationship database according to the query statement to obtain a search result is: performing semantic vector search in a preset spatial location relationship vector database according to the query statement vector to obtain a search result.
3. The method of claim 1, wherein, The step of responding to the user according to the search result comprises: determining whether the search result is unique; if yes, generating anaphora resolution data according to the search result to respond to the user; if no, providing the search results one by one to the user for confirmation, and after obtaining the user confirmation, generating anaphora resolution data according to the confirmed search result to respond to the user.
4. The method of claim 3, wherein, The step of providing the search results one by one to the user for confirmation, and after obtaining the user confirmation, generating anaphora resolution data according to the confirmed search result to respond to the user comprises: providing the search results one by one to the user for confirmation in a multi-modal manner; after obtaining the user confirmation, generating anaphora resolution data according to the confirmed search result to respond to the user.
5. The method of claim 1, wherein, The method further comprises: constructing a spatial location relationship database, the spatial location relationship database storing the name, visual description, spatial location relationship and function information of each spatial object.
6. The method of claim 2, wherein, The method further comprises: constructing a spatial location relationship database, the spatial location relationship database storing the name, visual description, spatial location relationship and function information of each spatial object; vectorizing the text information in the spatial location relationship database to obtain a text vector; storing the text vector into the database to construct a spatial location relationship vector database.
7. The method of claim 1, wherein, The spatial location relationship features include at least one of absolute position relationship and relative position relationship.
8. The method of claim 1, wherein, The visual description features include at least one of icon, shape, color and type. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to implement the spatial location relationship-based anaphora resolution method of any one of claims 1-8.
10. A computer-readable storage medium having stored thereon computer- executable instructions, wherein, The computer executable instructions are executed by a computer processor to implement the spatial location relationship-based anaphora resolution method of any one of claims 1-8.