Method for adjusting sound field of area based on inductive directional sound source, server and medium
By using inductive directional sound source technology, combining inductive signals and location descriptions, semantic segment markers are analyzed, audio content is pre-fetched, and the dwell tendency of target objects is analyzed to achieve a smooth transition of audio content. This solves the problem of inaccurate audio switching in existing technologies and improves the fluency of the auditory experience and the adaptability of the content.
Patent Information
- Application Number
- CN202611006115.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-08-25
AI Technical Summary
In existing technologies, a single audio segment is directly retrieved and played in a specific direction after the target object is detected by a sensor. This results in inaccurate switching of audio content, affecting the smoothness of the auditory experience, especially when the target object is moving or stationary, it cannot accurately match its behavior.
By linking sensing signals with location descriptions, the system analyzes semantic segmentation markers in the audio content, pre-fetches the next audio content, and generates dwell tendency based on the location fluctuations of the target object, thus achieving a smooth transition of audio content and ensuring synchronization between switching timing and playback progress.
It achieves seamless audio content delivery, improves the timing accuracy of audio push and the smoothness of the listening experience, adapts to changes in the behavior of the target audience, and maintains a coherent narrative output from directional sound sources.
Smart Images

Figure CN122640665A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more specifically, to a method for regional sound field adjustment based on an inductive directional sound source, a server, and a medium. Background Technology
[0002] In this invention, regionalized directional sound field technology is a technique for directionally pushing audio content to a target object within a target space based on the object's location. This can be applied to scenarios such as exhibition guidance and retail. Currently, sensors are typically used to detect the target object's location, and then a single audio segment pre-associated with that location is directly retrieved for directional playback. However, this location-based audio triggering method has limitations when handling continuously changing audio narrative needs. When the target object moves or remains within the area, the switching of audio content often relies on fixed location trigger conditions or simple playback end signals. This results in an inaccurate correspondence between content delivery and the target object's actual behavioral state, and the switching method between audio segments can easily affect the smoothness of the auditory experience. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide at least one method for regional sound field adjustment based on an inductive directional sound source, a server, and a medium.
[0004] According to one aspect of the present invention, a method for regional sound field adjustment based on an inductive directional sound source is provided, comprising: Receive the sensing signal corresponding to the target object generated by the sensor and the location description of the area where the target object is located, retrieve the first audio content associated with the location description according to the sensing signal and the location description, and control the directional sound source to output the first audio content to the area where the target object is located; During the output of the first audio content, the pre-embedded semantic segmentation markers within the first audio content are parsed to obtain the current output progress of the first audio content. Based on the semantic segmentation markers and the current output progress, the remaining output duration of the nearest semantic breakpoint is determined. Using the remaining output duration and position description as input, the second audio content is determined in the preset audio content library and prefetched into the output buffer. The system continuously collects the location description sequence of the target object generated by the sensor, forming a time-ordered set of location descriptions; it performs location fluctuation analysis on the time-ordered set of location descriptions to generate a description of the target object's tendency to stay in the current area, which characterizes the strength of the target object's tendency to transition from a wandering state to a stationary state. When the dwell tendency description reaches the preset tendency benchmark, and the output progress of the first audio content reaches any semantic breakpoint in the semantic segment division marker, the end of the output sampling sequence of the first audio content is continuously spliced with the beginning of the output sampling sequence of the second audio content in the output buffer to form an uninterrupted audio stream, and the directional sound source is controlled to output the second audio content. The second audio content is used as the new first audio content. The audio data that has not yet been output in the output buffer is retained as the current playback data source. The prefetch state of the output buffer is reset to idle. The step of parsing the semantic segment division markers pre-embedded in the first audio content during the output process of the first audio content is returned until the sensor stops generating the sensing signal of the target object.
[0005] According to another aspect of the present invention, a server is provided, comprising: a processor; and a memory, wherein the memory stores computer-readable code that, when executed by the processor, causes the processor to perform the method described above.
[0006] According to another aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method.
[0007] This invention triggers the output of first audio content by linking a sensing signal with a location description. During the output process, it analyzes semantic segmentation markers in real time and prefetches second audio content based on the current output progress and location description, allowing for advance preparation for subsequent audio switching. Simultaneously, it continuously collects the location description sequence of the target object and performs location fluctuation analysis to generate a dwell tendency description, characterizing the strength of the target object's tendency to transition from roaming to pausing. This ensures that the timing of audio content switching is determined by combining playback progress with behavioral state, avoiding abrupt switching before the target object has settled. When the dwell tendency reaches a benchmark and the connection is completed at the semantic breakpoint, a seamless audio stream is formed by continuously splicing the end of the sampled sequence with the beginning of the prefetched content, achieving a smooth transition between the two audio segments and eliminating auditory interruptions caused by switching gaps. After the switch is completed, the second audio content is used as the new first audio content and the prefetch state is reset, forming a continuous loop of semantic breakpoint switching and content prefetching. This allows the directional sound source to maintain a coherent narrative output throughout the target object's entire dwell time. The regional sound field content adapts and follows the target object's behavior changes, improving the timing accuracy of audio content delivery and the smoothness of the auditory experience. Attached Figure Description
[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the specification, serve to explain the technical solutions of the present invention.
[0009] Figure 1 This is a schematic diagram of an application scenario provided by the present invention; Figure 2 This is a flowchart illustrating a regional sound field adjustment method based on an inductive directional sound source provided by the present invention. Figure 3 This is a schematic diagram of the scheme logic of the method provided in the embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a server provided in an embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] To facilitate a clearer understanding of this invention, we will first introduce the application scenarios of the regional sound field adjustment method based on an inductive directional sound source, such as... Figure 1 As shown, the application scenario of this invention includes a server 10 and a sensor cluster. The sensor cluster may include one or more sensors; the number of sensors is not limited here. Figure 1 As shown, the sensor cluster may specifically include sensor 1, sensor 2, ..., sensor n; it can be understood that sensor 1, sensor 2, sensor 3, ..., sensor n can all be connected to server 10 via network so that each sensor can interact with server 10 via network connection.
[0012] It is understood that server 10 can refer to a server that executes the regional sound field adjustment method based on inductive directional sound source provided in the embodiments of the present invention, such as a single physical server, or a server cluster or distributed system consisting of at least two physical servers.
[0013] Further, please see Figure 2 and Figure 3 The regional sound field adjustment method based on an inductive directional sound source provided in this embodiment of the invention consists of... Figure 1 The method for regional sound field adjustment based on an inductive directional sound source is executed by server 10 and may include the following steps: Step S100: Receive the sensing signal corresponding to the target object generated by the sensor and the location description of the area where the target object is located, retrieve the first audio content associated with the location description according to the sensing signal and the location description, and control the directional sound source to output the first audio content to the area where the target object is located.
[0014] Sensors are devices or collaborative sensing networks composed of multiple types of sensors used to capture the state, identity attributes, or motion characteristics of target objects within a monitored space. Their sensing mechanisms can encompass one or more combinations of infrared pyroelectric sensing, ultrasonic echo ranging sensing, RFID tag sensing, LiDAR point cloud scanning sensing, or visual image analysis sensing. The sensing signal is the change in electrical signal or digitally encoded message generated by the sensor when it detects a target object entering or remaining within a specific spatial range, used to characterize the existence of the target object and its corresponding basic features. The target object is an individual that needs to be covered by the sound field service, such as visitors in an exhibition hall or consumers in a commercial complex. The area is the spatial range within which a directional sound source can effectively project a sound field and where the sensor has reliable sensing capabilities. Its boundaries can be defined by physical room partitions, infrared fences, the reading range of RFID reader antennas, or a virtual fence constructed by multi-line LiDAR. Location description is a digital representation of the spatial coordinates, sub-zone number, or trajectory of a target object within a region. Its data form can be a two-dimensional Cartesian coordinate pair based on a regional plan, a three-dimensional spatial coordinate set incorporating height information, a combination of polar azimuth and radial distance with a fixed reference point as the origin, or a sub-grid index encoding after pre-grid segmentation of the space. A directional sound source is a sound radiation device capable of projecting sound wave energy modulated by an audio signal in a narrow beam to a specific spatial location, achieving significant sound pressure attenuation in non-target directions. Typical implementations include directional speaker arrays composed of multiple electrodynamic full-range speaker units arranged in a spherical or planar phased array geometry, or parametric array acoustic transducers that use ultrasonic carrier waves to self-demodulate and generate audible sound in the air. The first audio content is an audio data unit with a static or dynamic binding relationship to the region. Its content can include audio explaining the background of exhibits, safety warning broadcasts, environmental audio adapted to the region's theme, or chapter-based audio programs with a complete narrative thread.
[0015] Step S200: During the output of the first audio content, the semantic segmentation markers pre-embedded in the first audio content are parsed to obtain the current output progress of the first audio content, and the remaining output duration of the nearest semantic breakpoint is determined based on the semantic segmentation markers and the current output progress; the remaining output duration and position description are used as input to determine the second audio content in the preset audio content library, and the second audio content is prefetched into the output buffer.
[0016] In one implementation, step S200 may specifically include the following steps S210 to S260: Step S210: Perform time-domain segmentation on the first audio content to obtain several overlapping short-time analysis panes. Extract the peak trend of the sound pressure level and the change pattern of the zero-crossing rate in each short-time analysis pane to form a zero-crossing rate-peak combination description sequence.
[0017] The short-time analysis pane is an audio signal analysis unit with a fixed analysis frame length and frame shift, extracted from the continuous sample point stream of the first audio content through a windowing function. It has a finite duration and some overlap between adjacent panes. The overlap between panes is designed to smooth out spectral leakage caused by rectangular window truncation and improve the temporal localization accuracy of rapidly changing events in non-stationary speech signals. The peak trend of the sound pressure level (SPL) is the trajectory of the maximum amplitude of the instantaneous SPL envelope of the audio signal within a single short-time analysis pane, depicting the distribution of instantaneous bursts of concentrated sound energy within the pane. The change in the zero-crossing rate (ZCR) is the microscopic evolution trend of the number of amplitude sign changes of audio sample points within the short-time analysis pane as time progresses within the pane. The level of the ZCR reveals the location of the main energy frequency band of the signal within that pane; a lower ZCR usually corresponds to voiced sounds or ambient hum with energy concentrated in the low-frequency region, while a higher ZCR often corresponds to unvoiced consonants or transient friction noise with energy dispersed in the mid-to-high frequency region. The zero-crossing rate-peak combined description sequence is a set of multi-dimensional feature vectors formed by jointly encoding the zero-crossing rate descriptors and sound pressure level peak descriptors obtained statistically in each short-time analysis pane according to the pane time sequence. This sequence preserves the evolution trajectory of the micro-acoustic features of the audio stream in the time domain.
[0018] Specifically, the pulse-code modulation data stream of the first audio content is divided into sliding frames according to preset frame length and frame shift parameters. A Hamming window with large sidelobe attenuation is selected as the window type to suppress spectral leakage. Within each short-time analysis pane, two levels of feature calculation are performed. The first level of calculation targets the peak trend of the sound pressure level. It searches for sampling points that satisfy the local maximum condition among the absolute value amplitudes of all sampling points in the pane. That is, the amplitude of the sampling point is greater than the amplitudes of its immediately preceding and immediately following sampling points. The searched local maxima points are arranged in chronological order, and their amplitude values are logarithmically scaled to match the human ear's perception of loudness, forming the peak amplitude trajectory of the pane. At the same time, the temporal distribution interval of these local maxima points is detected. If the interval shows an alternating pattern of dense and sparse intervals, this alternating rhythm is encoded as a peak density fluctuation description. The second-level calculation focuses on the zero-crossing rate variation pattern. Starting from the initial sampling point of the pane, it compares the positive and negative signs of the amplitudes of adjacent sampling points point by point. When a sign reversal occurs, the zero-crossing counter is incremented. For each fixed sub-window advance, the ratio of the accumulated zero-crossing count within the current sub-window to the total number of sampling points in the sub-window is recorded, forming a micro-variation curve of the zero-crossing rate within the pane. The mean, variance, slope trend, and local fluctuation amplitude statistics of this micro-variation curve are combined into a zero-crossing rate variation pattern vector. The peak amplitude trajectory, peak density fluctuation description, and zero-crossing rate variation pattern vector are tensor-concatenated with the zero-crossing rate variation pattern vector at the feature concatenation layer to generate the zero-crossing rate-peak combination description vector for this short-time analysis pane. All zero-crossing rate-peak combination description vectors generated for all panes are sorted according to the timestamp of the pane's initial sampling point, forming a zero-crossing rate-peak combination description sequence that strictly corresponds to the time sequence of the first audio content.
[0019] Step S220: In the zero-crossing rate-peak value combination description sequence, lock the silent interval segments where the zero-crossing rate is continuously lower than the silent zero-crossing threshold and the sound pressure level peak value is continuously lower than the silent energy threshold, extract the midpoint time of each silent interval segment, and generate a set of silent segmentation points.
[0020] The silence zero-crossing threshold is a zero-crossing rate decision boundary used to distinguish whether there are effective speech harmonic components in an audio signal. When the zero-crossing rate of a short-time analysis pane is lower than this threshold, it indicates that the energy of the signal in the pane is mainly concentrated in the extremely low-frequency region or has almost irregular frequency components, possessing the acoustic characteristics of a silent segment or background noise. The silence energy threshold is an amplitude decision boundary used to determine whether the signal sound pressure level has entered a weak level that can be ignored by the human ear. When the peak sound pressure level in a short-time analysis pane is continuously lower than this threshold, it indicates that there is no sound wave energy with perceptible intensity in that pane. A silent interval is an audio segment composed of multiple consecutive short-time analysis panes that simultaneously satisfies both the zero-crossing rate being lower than the silence zero-crossing threshold and the peak sound pressure level being lower than the silence energy threshold throughout the entire interval. This segment corresponds to intentional pauses, inter-sentence breathing gaps, or transitions between completely quiet scenes in the speech stream. The silence segmentation set is a set of timestamps used to represent the natural pauses in speech content, obtained by arranging the midpoint times of all identified silent intervals in ascending chronological order.
[0021] For example, sequentially traverse each combined description vector in the zero-crossing rate-peak value combined description sequence. Maintain a silent state machine, initially in a non-silent state. When a combined description vector with a zero-crossing rate below the silent zero-crossing threshold and a sound pressure level peak amplitude below the silent energy threshold is continuously scanned, the state machine switches from the non-silent state to the silent state, and records the start time of the first short-time analysis pane that meets the conditions as the silent start boundary. While the state machine remains in the silent state, it continues to receive subsequent combined description vectors. If a subsequent vector still simultaneously meets both conditions below the threshold, the silent duration counter is incremented. Once either condition is broken—that is, the zero-crossing rate rises above the silent zero-crossing threshold or the sound pressure level peak amplitude rises above the silent energy threshold—the state machine switches from the silent state back to the non-silent state, and records the start time of the current short-time analysis pane as the silent end boundary. Extract the midpoint time value between the silent start boundary and the silent end boundary, and write this midpoint time value as the representative moment of the silent interval segment into the silent segment candidate list. After traversing the entire zero-crossing rate-peak combination description sequence, redundancy removal is performed on all moments in the candidate list of silent split points. If the time interval between two midpoint moments is less than the duration of a short analysis pane, the one with the longer duration of the corresponding silent interval is retained, and finally, a set of silent split points is generated.
[0022] Step S230: Perform coherence analysis of harmonic components on the first audio content, detect the abrupt change positions of the harmonic structure between adjacent short-time analysis panes, extract the time of each harmonic structure abrupt change, and generate a set of harmonic inflection points.
[0023] In one implementation, step S230 may specifically include the following steps S231 to S236: Step S231: Perform multi-resolution spectral decomposition on each short-time analysis pane of the first audio content to generate a spectral linear array composed of several frequency components, and lock the set of candidate harmonic frequency positions in the spectral linear array based on the fundamental frequency prediction trajectory.
[0024] Multi-resolution spectral decomposition (MSD) is a spectral computing strategy that utilizes different time windows to perform multiple spectral analyses on the same audio frame, balancing the high frequency resolution requirements of low-frequency harmonics with the high temporal resolution requirements of high-frequency harmonics. A linear spectrum array, after MSD, stitches together the optimal spectral analysis results from different frequency bands into a dense spectral line covering the entire analysis band, with each frequency point corresponding to a complex amplitude value. The fundamental frequency prediction trajectory is the continuous change path of the fundamental frequency value over time estimated using a fundamental frequency tracking algorithm within a processed historical short-time analysis pane sequence. The candidate harmonic frequency position set is a series of discrete frequency positions derived from the fundamental frequency prediction value of the fundamental frequency prediction trajectory at the current pane, based on integer multiple harmonic relationships. These positions represent the spectral frequencies within the current pane that may carry harmonic energy.
[0025] For example, a short-time analysis pane sequence of the first audio content is received. For each short-time analysis pane, multi-resolution spectral decomposition is first performed. Specifically, a longer-duration Hanning window and a shorter-duration Hanning window are applied to the same short-time analysis pane. The data after windowing the long window is zero-padded to the same number of Fast Fourier Transform points and then the amplitude spectrum is calculated to obtain dense frequency sampling in the low-frequency region. The amplitude spectrum is also calculated for the data after windowing the short window to accurately capture rapid amplitude changes in the high-frequency region. The low-frequency part of the amplitude spectrum of the long window and the high-frequency part of the amplitude spectrum of the short window are smoothly spliced at the preset cross-frequency boundary using a cosine square weighted curve to form a linear spectral array covering the entire analysis frequency band with uniform frequency point spacing, where each frequency point is described by its amplitude and phase values. Simultaneously, the fundamental frequency prediction unit is invoked. This unit employs a cepstral method combined with dynamic programming smoothing, traversing the cepstral spectrum of each pane in the historical pane sequence. It searches for cepstral peaks within the cepstral interval corresponding to the fundamental frequency, and applies dynamic programming smoothing to the searched cepstral peak trajectories to eliminate jumps caused by misjudgments of fundamental frequency harmonics and half-frequency harmonics. This generates a fundamental frequency prediction trajectory spanning the historical panes and extending to the current pane. After obtaining the fundamental frequency prediction value corresponding to the current pane, using this fundamental frequency prediction value as the base frequency point, the integer multiple harmonic frequency positions within the first, second, and third harmonics of the fundamental frequency prediction value, up to the highest analysis frequency band, are marked on the spectral linear array. These frequency positions are collected as the candidate harmonic frequency position set for the current pane.
[0026] Step S232: Perform frequency component pairing on the spectral linear array of adjacent short-time analysis panes, establish a one-to-one correspondence between the position of each candidate harmonic frequency in the preceding pane and the component with the smallest frequency offset in the following pane, and form a cross-window harmonic pairing link.
[0027] Specifically, temporally adjacent preceding and following short-time analysis panes can be used as processing pairs. Each candidate harmonic frequency position in the preceding pane's candidate harmonic frequency position set is traversed. An asymmetric search range is set on the frequency axis centered on that frequency position. The upper and lower bounds of this search range are set according to the order of the current harmonic; lower-order harmonics use a narrower search range to maintain the rigor of frequency tracking, while higher-order harmonics use a slightly wider search range to accommodate possible frequency drift. In the spectral linear array of the following pane, only actual frequency components within this search range are selected as candidate matching targets. The absolute value of the frequency difference between each candidate matching target's frequency and the reference harmonic frequency position is calculated, and the actual frequency component with the smallest absolute difference is selected as the paired component. If no actual frequency component exists within the search range, the cross-window pairing is marked as missing. The identification information of the harmonic frequency position in the preceding pane, the frequency and amplitude information of the paired component in the following pane, and the frequency offset between them are recorded together as a pairing record. After all candidate harmonic frequency positions in the preceding pane are paired, the frequency components of the following panes in each pairing record are used as the preceding references for the next level of pairing. This process is iterated sequentially according to the pane progression direction, concatenating the paired components belonging to the same harmonic source in each pane along the time axis to form several cross-window harmonic pairing links. Each cross-window harmonic pairing link contains at most one harmonic component at each short-time analysis pane position, and the link index remains constant throughout the entire audio stream span, enabling independent tracking of each harmonic sequence.
[0028] Step S233: Trace the amplitude evolution curve and frequency offset curve of each harmonic component one by one along the cross-window harmonic pairing link, detect the moment when the slope reverses in the amplitude evolution curve and the moment when the direction reverses in the frequency offset curve, and mark them as amplitude anomalies and frequency anomalies, respectively.
[0029] The amplitude evolution curve is the trajectory of amplitude change over time, formed by extracting the amplitude values of the corresponding harmonic components in each short-time analysis pane along the time axis of a cross-window harmonic pairing link. The frequency offset curve is the trajectory of frequency deviation change, formed by sorting the frequency values of the harmonic components in each pane along the same link by time, based on integer multiples of the reference fundamental frequency. Slope reversal refers to the inversion of the first derivative sign of the amplitude evolution curve around a certain moment, i.e., the turning point where the curve changes from monotonically rising to monotonically falling, or vice versa. Directional reversal refers to the turning point where the frequency offset curve changes from a continuous shift towards higher frequencies to a shift towards lower frequencies, or vice versa. An amplitude anomaly is the specific moment when a slope reversal is detected on the amplitude evolution curve. A frequency anomaly is the specific moment when a directional reversal is detected on the frequency offset curve.
[0030] For example, each cross-window harmonic pairing link is extracted, and two time-varying sequences are constructed for each link: an amplitude evolution sequence and a frequency offset sequence. The amplitude evolution sequence is generated by reading amplitude values from the paired component records of each pane, using the short-time analysis pane number spanned by the link as the horizontal axis, forming a discrete amplitude sampling sequence. A moving average filter is applied to this sequence to remove minor jitter fluctuations, resulting in a smooth amplitude evolution curve. The frequency offset sequence is generated by calculating the offset between the actual frequency of each paired component in each pane and the ideal harmonic frequency obtained by multiplying the current pane's fundamental frequency estimate by the harmonic order. These offsets are arranged by pane number to form a frequency offset sequence, which is also filtered using a moving average filter to obtain the frequency offset curve. A first-order difference calculation is performed on the amplitude evolution curve to obtain an amplitude slope sequence. This sequence is scanned point by point, and when two consecutive difference values are found to have opposite signs, that position is identified as a candidate point for slope inversion. The frequency offset curve is analyzed to determine the offset direction. This determination is based on the sign of the increments of frequency offset values across multiple consecutive panes. A positive increment indicates a shift towards higher frequencies, while a negative increment indicates a shift towards lower frequencies. When a reversal in the offset direction is detected, that location is identified as a candidate point for directional reversal. The time coordinates of the confirmed slope reversal candidate points on the amplitude evolution curve are marked as amplitude anomalies, and the time coordinates of the confirmed directional reversal candidate points on the frequency offset curve are marked as frequency anomalies. Each anomaly is then appended with its corresponding harmonic order and link identifier.
[0031] Step S234: In the same cross-window harmonic pairing link, when the time interval between the amplitude anomaly point and the frequency anomaly point is less than the preset synchronization anomaly judgment interval, the time position is taken as the instability position of the link; otherwise, the independent instability positions corresponding to the amplitude anomaly point and the frequency anomaly point are retained respectively.
[0032] In one implementation, step S234 may specifically include the following steps S2341 to S2346: Step S2341: For each cross-window harmonic pairing link, traverse its amplitude evolution curve along the time axis. When an inflection point is detected where the slope of the amplitude evolution curve changes from positive to negative or from negative to positive, record the time corresponding to the inflection point as a candidate point for amplitude anomaly.
[0033] For example, the smooth amplitude evolution curve of a cross-window harmonic pairing link is scanned, the amplitude difference between adjacent sampling points is calculated, and a first-order difference sequence is constructed. The locations where the sign changes in the first-order difference sequence are identified, i.e., the previous difference is positive and the next difference is negative, or vice versa. At the sign change location, it is confirmed that there are a sufficient number of valid amplitude sampling points on both sides of the location to avoid false inflection points caused by boundary effects. After confirmation, the time coordinate corresponding to the location is recorded as a candidate point for amplitude anomalies, and this candidate point is bound and stored with the cross-window harmonic pairing link identifier and link order.
[0034] Step S2342: Screen the candidate points of amplitude anomaly, calculate the steepness of the slope change of the amplitude evolution curve before and after the point, and retain only the candidate points of amplitude anomaly whose slope change exceeds the preset steepness threshold as amplitude anomaly points.
[0035] The steepness of the slope change measures the abruptness of the change in the slope of the amplitude evolution curve at the inflection point, from a positive maximum to a negative maximum or vice versa. This measure can be assessed by calculating the difference between the mean of the first-order differences of several sampling points before the inflection point and the mean of the first-order differences of several sampling points after the inflection point. A larger difference indicates a more abrupt amplitude change. The preset steepness threshold is a criterion used to distinguish between rapid amplitude changes truly caused by switching vocalization actions and slow amplitude changes caused by normal volume fluctuations.
[0036] Specifically, for each candidate amplitude anomaly point, the first-order difference values of the amplitude evolution curves of multiple consecutive sampling points in the preceding time direction are extracted, and the arithmetic mean of these preceding difference values is calculated as the preceding average slope. Then, the first-order difference values of multiple consecutive sampling points in the following time direction are extracted, and the arithmetic mean of these following difference values is calculated as the following average slope. The absolute difference between the preceding and following average slopes is calculated, and this absolute difference is used as a measure of the steepness of the slope change at the candidate point. This measure is compared with a preset steepness threshold. If the measure is greater than the preset steepness threshold, the amplitude change at the candidate point is determined to be sufficiently drastic, and it is confirmed as an amplitude anomaly point. If the measure is not greater than the preset steepness threshold, the candidate point is determined to be a normal volume fluctuation inflection point and is discarded. All amplitude anomaly points that pass the screening are stored in an amplitude anomaly point set, retaining their time coordinates and the index of the cross-window harmonic pairing link to which they belong.
[0037] Step S2343: Traverse the frequency offset curve along the time axis of the same cross-window harmonic pairing link. When a switching point is detected where the offset direction of the frequency offset curve changes from offset to high frequency to offset to low frequency or from offset to offset to high frequency, record the time corresponding to the switching point as a candidate point for frequency anomaly.
[0038] The direction of the frequency offset curve refers to whether the deviation of the actual frequency of a harmonic component from its ideal harmonic frequency position continues to increase or decrease within a certain time window. The direction switching point is the instant when this offset trend reverses, such as the turning point from a continuous drift towards higher frequencies to a return to lower frequencies. Frequency anomaly candidate points are the moments corresponding to all initially identified direction switching points, without yet being verified by the degree of abrupt change in the offset rate.
[0039] For example, a frequency offset sequence is extracted along the time axis of the same cross-window harmonic pairing link. A direction determination window is set and slides along the frequency offset sequence from the starting point. The linear fitting slope of the frequency offset value is calculated within the window. If the slope is consistently positive and exceeds the shallow fluctuation tolerance range, the offset direction within the window is determined to be a shift towards higher frequencies; if the slope is consistently negative and the absolute value exceeds the shallow fluctuation tolerance range, it is determined to be a shift towards lower frequencies. The offset direction determination results of adjacent direction determination windows are compared one by one. When a difference in the direction determination results of adjacent windows is detected, it is determined that a direction switch has occurred, and the direction switch point is marked at the boundary time of the two windows. The time coordinate of the direction switch point is recorded as a candidate point for frequency anomaly.
[0040] Step S2344: Confirm the candidate points of frequency anomalies, calculate the change in the offset rate of the frequency offset curve before and after the point, and retain only the candidate points of frequency anomalies whose change in offset rate exceeds the preset rate change threshold as frequency anomaly points.
[0041] For each candidate frequency anomaly, a first-order difference sequence of several consecutive preceding sampling points is extracted from the frequency offset sequence. The absolute average of this preceding difference sequence is calculated as the forward offset rate representation. Then, a first-order difference sequence of several consecutive following sampling points is extracted, and the absolute average of this following difference sequence is calculated as the backward offset rate representation. The absolute difference between the forward and backward offset rate representations is calculated as the offset rate change of the candidate point. This offset rate change is compared with a preset rate mutation threshold. If the offset rate change is greater than the preset rate mutation threshold, the candidate frequency anomaly is confirmed as a frequency anomaly; if the offset rate change is not greater than the preset rate mutation threshold, the candidate point is excluded. All confirmed frequency anomalies are stored in a frequency anomaly set.
[0042] Step S2345: Arrange the amplitude anomaly points and frequency anomaly points in the same cross-window harmonic pairing link in chronological order. When the time span between adjacent amplitude anomaly points and frequency anomaly points is less than the preset synchronization anomaly judgment interval, take the midpoint of their time as the instability position.
[0043] For example, for each cross-window harmonic pairing link, extract the amplitude anomaly list and frequency anomaly list belonging to that link, merge the two lists, and sort them in ascending order of time. Traverse the sorted anomaly sequence. For each amplitude anomaly, search the sequence for the nearest frequency anomaly both forward and backward, and calculate the absolute value of the time difference between the amplitude anomaly and the frequency anomaly. For each frequency anomaly, similarly search the sequence for the nearest amplitude anomaly and calculate the absolute value of the time difference. When the absolute value of the time difference between two adjacent anomaly points is less than the preset synchronization anomaly judgment interval, record these two anomalies as a synchronization anomaly pair. Calculate the arithmetic mean of the two times in the synchronization anomaly pair, mark the average value as the instability position, and remove the two anomalies that make up the synchronization anomaly pair from their respective lists to avoid duplicate pairing. If there are no frequency anomalies that meet the synchronization conditions before or after an amplitude anomaly, or no amplitude anomalies that meet the synchronization conditions before or after a frequency anomaly, these anomalies do not participate in pairing and are left for processing in subsequent steps.
[0044] Step S2346: When there are no frequency anomalies before or after the amplitude anomaly point with a time span smaller than the preset synchronization anomaly judgment interval, the amplitude anomaly point is regarded as an independent instability location; when there are no amplitude anomalies before or after the frequency anomaly point with a time span smaller than the preset synchronization anomaly judgment interval, the frequency anomaly point is regarded as an independent instability location.
[0045] For example, after completing the merging of synchronization anomalies, the remaining unpaired anomalies in the amplitude and frequency anomaly lists are re-examined. For each remaining amplitude anomaly, its originally recorded time coordinate is directly added to the instability location set as the instability location, and its type is marked as amplitude-independent instability. For each remaining frequency anomaly, its time coordinate is similarly added to the instability location set as the instability location, and its type is marked as frequency-independent instability. After this step, all detected anomalies on each cross-window harmonic pairing link are converted into instability location records in a uniform format.
[0046] Step S235: Count the number of links with unstable positions in each cross-window harmonic pairing link at the boundary of the same short-time analysis pane. When the ratio of the number of such links to the total number of cross-window harmonic pairing links exceeds the preset instability ratio threshold, it is determined that a harmonic structure mutation has occurred at the boundary of the current short-time analysis pane.
[0047] The boundary of the same short-time analysis pane refers to the timeline between two adjacent short-time analysis panes. The two sides of this boundary line represent the tail and head sampling points of the preceding and following panes, respectively, and are the basic analytical unit for assessing the temporal granularity of changes in the overall harmonic structure. The total number of harmonic pairs across the window is the total number of harmonic links that maintain a pairing relationship between the preceding and following pane pairs in the current analysis. The instability ratio threshold is the minimum percentage of links with instability necessary to determine a complete harmonic structure abrupt change event. Setting this threshold can exclude sporadic instabilities caused by local noise interference in a few links. A harmonic structure abrupt change refers to a reorganization of the overall harmonic structure across the pane sufficient to alter the timbre or vocal type, typically corresponding to phoneme boundaries, syllable boundaries, or vocal type transitions.
[0048] For example, the boundary between adjacent short-time analysis panes is used as the analysis unit, and each pane boundary is scanned one by one. For the currently processed pane boundary, all cross-window harmonic pairing links crossing the boundary are summarized, and the number of links to which the instability position belongs within one pane time range from the boundary in time is counted as the number of unstable links. At the same time, the total number of all cross-window harmonic pairing links with pairing records at the pane boundary is counted as the total number of links. The ratio of the number of unstable links to the total number of links is calculated, and this ratio is compared with a preset instability ratio threshold. When the ratio exceeds the preset instability ratio threshold, it is determined that a harmonic structure abrupt change has occurred at the pane boundary, and the time of the boundary is marked as a candidate time of harmonic structure abrupt change. If the ratio does not exceed the threshold, the pane boundary is regarded as a harmonic structure smooth transition boundary and is not recorded.
[0049] Step S236: Collect all short-time analysis pane boundary moments that are determined to have caused harmonic structural abrupt changes, arrange them in chronological order and remove consecutively repeating adjacent boundary moments to generate a set of harmonic inflection points.
[0050] Specifically, all window boundary times marked as harmonic structure abrupt changes in step S235 are collected, and these times are arranged in ascending chronological order to form an initial time list. This initial time list is traversed, maintaining a sliding comparison window. If the time difference between the current boundary time and the most recent retained harmonic inflection point time is less than the duration of a short-time analysis window, it is considered a continuous repetition, and the current boundary time is skipped and not retained. If the time difference between the current boundary time and the most recent retained harmonic inflection point time is not less than the duration of a short-time analysis window, the current boundary time is added as a new harmonic inflection point to the result list. After traversal, all times stored in chronological order in the result list constitute the harmonic inflection point set.
[0051] Step S240: Perform time-axis union fusion of the set of silent breakpoints and the set of harmonic inflection points. When the time interval between the silent breakpoint and the harmonic inflection point is less than the preset fusion tolerance interval, one of them is removed; otherwise, both are retained to form a coarsely selected breakpoint set.
[0052] Time axis union fusion is the process of placing two independent sets of breakpoint timestamps from the silence detection and harmonic analysis pathways onto a unified time axis and merging and deduplicating them in chronological order. The preset fusion tolerance interval is a criterion used to determine the temporal proximity of silence breakpoints and harmonic inflection points to the same semantic pause position. If the time difference between the two is extremely small, they can be considered repeated discoveries of different acoustic representations of the same natural breakpoint. The coarse breakpoint set is a preliminary selection of candidate semantic breakpoints obtained after time axis union fusion and proximity point deduplication; each time point in this set represents a possible semantic segment transition position.
[0053] Specifically, all moments from the silence segmentation point set and all moments from the harmonic inflection point set are extracted into a single list of time points to be processed. Each time point is marked with a source tag to distinguish whether it originated from silence detection or harmonic analysis. The list of time points to be processed is then sorted in ascending order of time. After sorting, a two-pointer traverses the list, with each pointer tracking the most recently unprocessed silence segmentation point and harmonic inflection point. When the absolute value of the time difference between a silence segmentation point and a harmonic inflection point is less than a preset fusion tolerance interval, the point is removed according to a retention strategy. This strategy can be set to prioritize retaining harmonic inflection points, as they are usually more closely related to changes in vocalization in time, or to retain the point with the higher original confidence level. If the absolute value of the time difference is not less than the preset fusion tolerance interval, both points are retained. After traversal and removal, all remaining time points are rearranged in ascending order of time to form a coarsely selected breakpoint set.
[0054] Step S250: Perform contextual semantic coherence filtering on the coarsely selected set of breakpoints, remove isolated breakpoints that appear in the middle of a coherent speech segment, and obtain semantic segment division markers.
[0055] For example, each breakpoint moment in the coarsely selected breakpoint set is traversed. For each breakpoint moment, an audio segment is extracted, extending forward and backward by a predetermined context observation duration centered on that breakpoint. A speech activity detector is run on this audio segment, employing a decision logic based on a dual threshold of long-term spectral flatness and short-term energy to label the audio frame as either a spoken frame or a silent frame. The distribution of consecutive spoken frames before and after the breakpoint is statistically analyzed. If a long, continuous spoken segment exists before the breakpoint, and a similarly long, continuous spoken segment follows immediately after the breakpoint, and no other coarsely selected breakpoints exist within this consecutive spoken area, then the breakpoint is considered an isolated breakpoint and is discarded. If there is a significant silent gap before or after the breakpoint, or if there are significant differences in energy and spectral characteristics between the preceding and following spoken segments, indicating that they belong to different vocal units, then the breakpoint is retained. The breakpoint moments retained after contextual semantic coherence filtering are arranged in chronological order to constitute the final semantic segmentation markers.
[0056] Step S260: Based on the remaining output duration and location description, query the preset audio content library for multiple versions of narrative audio corresponding to the location description. Based on the narrative branching link clues recorded in the first audio content, select one version from the multiple versions of narrative audio as the second audio content and store it in the output cache.
[0057] For example, the system receives the remaining output duration calculated in step S200 and the latest location description from the sensor. Using the region identifier parsed from the location description as the query key, it retrieves a list of multiple versions of narrative audio corresponding to the region from the relational index table of the preset audio content library. Each audio record in the list contains an audio file identifier, duration, version tag, and content summary. Simultaneously, it reads the metadata of the narrative bifurcation link clue embedded in the first audio content. This metadata describes a decision table at the current semantic segment division marker point. The decision table takes the remaining output duration and the historical location residency characteristics of the target object as input conditions. The current remaining output duration and location description are input into the decision table, which outputs an audio version tag through table lookup matching. Based on the output version tag, the audio version matching the version tag is selected from the list of multiple versions of narrative audio as the second audio content. A direct memory access read request is then initiated to the storage subsystem to move the pulse-code modulation data block of the second audio content from the disk array or solid-state storage volume to the memory-mapped output buffer area, and pushes the write pointer of the output buffer to the end of the data, completing the prefetch operation.
[0058] Step S300: Continuously collect the position description sequence of the target object generated by the sensor to form a time-ordered set of position descriptions; perform position fluctuation analysis on the time-ordered set of position descriptions to generate a description of the target object's tendency to stay in the current area. The description of the tendency to stay represents the strength of the target object's tendency to change from a wandering state to a stationary state.
[0059] In one implementation, step S300 may specifically include the following steps S310 to S360: Step S310: Divide the location description time sorting set into multiple backtracking windows in chronological order. Each backtracking window contains several continuously collected location descriptions. Adjacent backtracking windows maintain a sliding progressive relationship of window length translation step.
[0060] For example, maintain a first-in, first-out (FIFO) time-ordered set of location descriptions, pushing new location descriptions to the end of the set as they arrive. Starting from the earliest time in the set, according to a preset window length and translation step, extract the first backtracking window on the timeline, encompassing all location descriptions within the window. Subsequently, slide the window's start time forward by one translation step to extract the second backtracking window, overlapping with the first in time. Continue generating the backtracking window sequence until the window's end time touches the latest time in the time-ordered set of location descriptions, thus dividing the entire time-ordered set of location descriptions into a series of fixed-length, incrementally sliding backtracking windows.
[0061] Step S320: In each backtracking window, calculate the spatial distribution dispersion of the location description to obtain the spatial dispersion pattern corresponding to the backtracking window, and compare the spatial dispersion patterns of all backtracking windows across windows to generate a spatial dispersion contraction trend indicator.
[0062] Spatial dispersion is a statistical measure of the degree of dispersion of all locations within a backtracking window in two-dimensional or three-dimensional space. Higher dispersion indicates a larger activity range and less concentrated location of the target object within that backtracking window; lower dispersion indicates a more concentrated location of the target object. Spatial dispersion pattern is a single-window descriptor that describes the dynamic changes in the target object's activity range over time by binding spatial dispersion to the backtracking window timestamp. Spatial dispersion contraction trend indication identifies a trend marker where spatial dispersion shows a continuous downward trend by comparing spatial dispersion pattern sequences from multiple consecutive backtracking windows. The emergence of this marker indicates that the target object's spatial activity range is gradually converging.
[0063] For each backtracking window, the two-dimensional or three-dimensional spatial coordinates of all location descriptions within the window are extracted to construct a location covariance matrix. Eigenvalue decomposition is performed on this covariance matrix to obtain the largest eigenvalue and the second largest orthogonal eigenvalue. The square root of the largest eigenvalue is used as the principal component of the spatial dispersion, or the geometric mean of the two larger eigenvalues is used as a comprehensive measure of the spatial dispersion, reflecting the overall diffusion radius of the location descriptions deviating from their spatial mean in various directions. This spatial dispersion, along with the window's start and end timestamps, is packaged into a spatial dispersion status record for that backtracking window. After calculating the spatial dispersion status for all backtracking windows, the spatial dispersion values of adjacent backtracking windows are compared one by one along the time progression direction. When a continuous decrease in spatial dispersion is detected across multiple consecutive backtracking windows, and the cumulative decrease exceeds a preset contraction significance threshold, a spatial dispersion contraction trend indicator is generated. A true value for this indicator indicates a significant convergence of the spatial activity range.
[0064] Step S330: In each backtracking window, extract the cumulative turning angle of the displacement segment formed by adjacent position descriptions to obtain the path detour description corresponding to the backtracking window, and perform cross-window comparison of the path detour descriptions of all backtracking windows to generate a path detour upward trend indicator.
[0065] In one implementation, step S330 may specifically include the following steps S331 to S336: Step S331: Obtain the sequence of position descriptions arranged in chronological order within the current backtracking window, connect each position description with its immediately following position description to form a displacement segment, and record each displacement segment as a displacement segment vector containing direction angle and length.
[0066] The orientation angle is the directed angle of the displacement segment vector relative to a reference direction in a preset reference coordinate system. The reference direction can be set to the positive horizontal direction of the region planar map, and the value range of the orientation angle covers the entire circumference angle. The length refers to the Euclidean distance of the displacement segment vector, representing the straight-line distance the target object moves between two adjacent sampling points. The displacement segment vector is a composite data unit that simultaneously contains both orientation angle and length attributes, used to fully characterize the movement gait between adjacent sampling points.
[0067] For example, extract the sequence of position descriptions arranged in ascending time order within the current backtracking window. Starting with the first position description in the sequence, subtract the coordinates of each position description from the coordinates of its immediately following position description to obtain the horizontal and vertical components of the displacement increment. Using the horizontal and vertical components, calculate the orientation angle by looking up a pre-stored arctangent lookup table in memory. Simultaneously, calculate the square root of the sum of the squares of the horizontal and vertical components to obtain the length. Encapsulate the orientation angle and length into a displacement segment vector structure, and form a displacement segment vector sequence according to the chronological order of the position descriptions.
[0068] Step S332: Calculate the direction deflection angle of each displacement segment vector relative to its immediate preceding displacement segment vector, map the direction deflection angle to a steering angle with positive and negative signs, assign a positive sign to clockwise deflection and a negative sign to counterclockwise deflection, and generate a steering angle sign sequence.
[0069] The direction deflection angle is the angular difference between the direction angle of the subsequent displacement segment vector and the direction angle of the previous displacement segment vector. This difference needs to be normalized to a range of -180 degrees to +180 degrees to eliminate the multi-valued nature of the angle. The steering angle is the angle value after assigning a sign to the direction deflection angle. The sign definition rule is that clockwise deflection corresponds to a positive sign, and counterclockwise deflection corresponds to a negative sign. The steering angle sign sequence is a sequence of signed angle values formed by arranging all steering angles in the backtracking window in the order of their appearance.
[0070] Step S333: Divide the steering angle symbol sequence into consecutive segments with the same sign, count the absolute cumulative amount of steering angle in each consecutive segment with the same sign and the time duration of the segment, and extract the switching frequency of alternating segments with the same sign.
[0071] A continuous segment with the same sign is a subsequence within a steering angle sign sequence where all steering angles within this subsequence are either all positive or all all negative, representing that the target object continuously turns in the same direction during that time period. The absolute cumulative amount is the sum of the absolute values of all steering angles within a continuous segment with the same sign, reflecting the total angle turned by the target object within that segment. The time duration is the length of time spanned by a continuous segment with the same sign, from the appearance of the first steering angle to the end of the last. The switching frequency is the total number of times the steering angle sign sequence alternates between positive and negative, corresponding to the number of times the target object switches from clockwise to counterclockwise or vice versa.
[0072] For example, a sequence of steering angle symbols is scanned. At the start of the scan, the symbol of the first non-zero steering angle in the sequence is read as the symbol of the current segment with the same number. The start timestamp of the segment is recorded, and the absolute cumulative value of the steering angle for that segment is initialized to the absolute value of the first steering angle. Subsequent steering angles are read sequentially. If the current steering angle symbol is the same as the symbol of the current segment with the same number, the absolute value of the steering angle is added to the absolute cumulative value, and the end timestamp of the segment is updated. If the current steering angle symbol is opposite or the steering angle is 0, the current segment with the same number ends. The symbol, absolute cumulative value, and duration of that segment are recorded in the segment table with the same number. Simultaneously, a new segment with the same number is created starting with the current steering angle, and the switching frequency counter is incremented. After scanning the entire sequence of steering angle symbols, the list of segments with the same number and the switching frequency are output.
[0073] Step S334: The ratio of the absolute cumulative amount of each consecutive segment with the same number to the duration of the segment is taken as the turning density of the segment. Then, the turning density of all consecutive segments with the same number in the current backtracking window is statistically aggregated at the window level to obtain the representative value of the window turning density. At the same time, the switching frequency of alternating segments with the same number is taken as the window switching frequency. The representative value of the window turning density and the window switching frequency are normalized and then merged to generate the path detour description of the backtracking window.
[0074] For example, the list of segments with the same number in the current backtracking window is traversed. For each consecutive segment with the same number, its absolute cumulative value is divided by the time duration to obtain the turning density of that segment. Using the time duration of each segment as a weight, a weighted average of the turning densities of all segments is calculated to obtain a representative value for the window's turning density. The switching frequency is directly used as the window switching frequency. Then, both values are normalized, with the normalization reference base values taken from the upper limits of turning density and switching frequency when a typical target object moves within the area, based on offline statistics. The normalized representative value for window turning density and the normalized window switching frequency are linearly weighted and summed according to a preset fusion weight to obtain a path detour description value. This path detour description value is recorded in the description structure of the backtracking window.
[0075] Step S335: Compare the path detour descriptions of adjacent backtracking windows one by one along the time progression direction. When the path detour description shows a monotonically increasing trend and continues for at least a preset number of consecutive backtracking windows, it is determined that an upward trend in path detour has been detected.
[0076] For example, maintain an upward trend counter with an initial count of 0. Starting from the second backtracking window, compare its path detour description value with that of the preceding backtracking window. If the current window value is greater than or equal to the previous window value, increment the upward trend counter by 1; if the current window value is less than the previous window value, reset the upward trend counter to zero. Check the current value of the upward trend counter. When the count value reaches or exceeds a preset number of consecutive backtracking windows, it is determined that an upward trend in path detour has been detected, and this determination result, along with the timestamp of the current window, is output as an upward trend indication of path detour.
[0077] Step S336: After determining that an upward trend in path detourness has been detected, further verify whether the turning angle symbol sequence of each backtracking window in the increasing interval shows a synchronous increase in the switching frequency of the same number segment. When the switching frequency of the same number segment increases synchronously with the path detourness description, it is confirmed as an upward trend indication of path detourness and output.
[0078] For example, in step S335, the module extracts the continuously increasing backtracking window intervals identified in step S335, and obtains the window switching frequency sequence for each backtracking window within that interval. The first-order difference of this window switching frequency sequence is calculated, and the sign of these first-order differences is checked. If the window switching frequency sequence also satisfies the monotonically increasing condition that the later value is not less than the earlier value within the increasing interval, the module confirms that the switching frequency of the same-number segment and the path detour description have a synchronous increasing relationship, marks the previously generated path detour upward trend indicator as valid, and officially outputs this indicator. If the window switching frequency sequence does not show a monotonically increasing trend, the module cancels the output of the path detour upward trend indicator and continues normal monitoring.
[0079] Step S340: For each backtracking window, calculate the ratio of the acquisition time span of all location descriptions within the window to the number of location descriptions that appear, obtain the location update density corresponding to the backtracking window, and compare the location update density of all backtracking windows across windows to generate an indicator of the trend of update density change.
[0080] In one implementation, step S340 may specifically include the following steps S341 to S346: Step S341: Within each backtracking window, extract the time interval between the acquisition times described by each adjacent location to form a location update time interval sequence, and perform interval distribution morphology analysis on the location update time interval sequence to distinguish between dense interval clusters and sparse interval clusters.
[0081] The location update time interval sequence is a sequence composed of the timestamp differences described by every two adjacent positions within the backtracking window, arranged chronologically. Interval distribution pattern analysis involves interpreting the statistical distribution characteristics of the time interval sequence to identify which time intervals belong to relatively short, densely distributed regions and which belong to relatively long, sparsely distributed regions. Dense interval clusters are groups of time intervals that are relatively close in time and have small numerical values. Sparse interval clusters are groups of time intervals with large numerical values.
[0082] Specifically, the location description sequence within the backtracking window is traversed. Starting from the second item, the timestamp of each item is subtracted from the timestamp of the previous item to generate a location update time interval sequence. Cluster analysis is performed on all time interval values in the location update time interval sequence using a distance-based partitioning algorithm, such as K-centroid clustering, to divide the sample points into two clusters. The average time interval for each cluster is calculated, and the cluster with the smaller average time interval is labeled as a dense cluster, while the cluster with the larger average time interval is labeled as a sparse cluster.
[0083] Step S342: Calculate the average time interval within the densely spaced clusters and the sparsely spaced clusters respectively, and compare the average time interval of the densely spaced clusters with the average time interval of the sparsely spaced clusters to generate the inter-cluster interval span ratio.
[0084] The intra-cluster average time interval is the arithmetic mean of all time interval values within a densely spaced or sparsely spaced cluster, representing the typical update interval for dense and sparse periods, respectively. Span comparison involves dividing or rationing the average time interval of the densely spaced cluster to that of the sparsely spaced cluster to measure the degree of separation between the two clusters. The inter-cluster interval span ratio is the ratio of the average time interval of the densely spaced cluster to that of the sparsely spaced cluster; a larger ratio indicates a more significant difference between the two update density states.
[0085] Specifically, the arithmetic mean of all time interval values within both densely spaced and sparsely spaced clusters is calculated to obtain the average interval of the dense clusters and the average interval of the sparse clusters. The average interval of the sparse clusters is divided by the average interval of the dense clusters to obtain the inter-cluster interval span ratio, which is always not less than 1. The inter-cluster interval span ratio is recorded as the sparse-dense separation index for this backtracking window.
[0086] Step S343: Divide the position update time interval sequence into several continuous segments according to time sequence, count the fluctuation degree of the time interval within each continuous segment, detect the transmission direction of the fluctuation degree between continuous segments, and generate the interval fluctuation transmission trend.
[0087] Specifically, the location update time interval sequence is divided into multiple continuous and non-overlapping segments. For each segment, the variance of all time intervals within that segment is calculated as the fluctuation level of that segment. These fluctuation levels are arranged in chronological order to form a fluctuation sequence. A trend test is performed on this fluctuation sequence by calculating its rank correlation coefficient to determine the direction of its monotonic trend. If the rank correlation coefficient is positive and passes the significance test, a transmission trend description indicating increasing interval fluctuation is generated; if the rank correlation coefficient is negative and passes the significance test, a transmission trend description indicating decreasing interval fluctuation is generated; if the rank correlation coefficient is not significant, it is determined that there is no obvious transmission trend.
[0088] Step S344: By combining the inter-cluster interval span ratio and the interval fluctuation transmission trend, determine whether the position update in the current backtracking window shows an intermittent contraction trend. When the inter-cluster interval span ratio is higher than the preset inter-cluster convergence threshold and the interval fluctuation transmission trend points to the narrowing of fluctuation, generate a contraction determination mark.
[0089] For example, the inter-cluster interval span ratio of the current backtracking window is compared with a preset inter-cluster convergence threshold, and the interval fluctuation transmission trend generated in step S343 is read. When the inter-cluster interval span ratio is greater than the preset inter-cluster convergence threshold, and the interval fluctuation transmission trend points to a narrowing of fluctuations, it is determined that the position update within the current backtracking window exhibits an intermittent contraction trend, and the contraction determination flag of the backtracking window is set to true. If any of these conditions are not met, the contraction determination flag is set to false.
[0090] Step S345: Traverse all backtracking windows one by one along the sliding direction, and construct the density description of each backtracking window by combining the shrinkage determination identifier of each backtracking window with its inter-cluster interval span ratio. Then, transfer and compare the density descriptions of adjacent backtracking windows to detect the frequency and continuity of the shrinkage determination identifier.
[0091] The density description is a snapshot of the backtracking window's position update state, composed of the contraction determination flag and the inter-cluster span ratio. It comprehensively records the sparse-density state and contraction evolution signs within each window. Transition alignment compares the density descriptions of adjacent backtracking windows to observe the evolutionary pattern of the contraction determination flag changing from false to true or remaining true. Occurrence frequency refers to the proportion of backtracking windows with a true contraction determination flag within a sliding window interval. Continuity refers to whether backtracking windows with a true contraction determination flag form continuous, uninterrupted segments.
[0092] Step S346: When the number of backtracking windows with consecutive contraction determination marks exceeds the preset threshold for the number of consecutive windows, and the density of position updates of each backtracking window in the consecutive segment shows a monotonically decreasing trend, an update density change trend indicator is generated and output. The update density change trend indicator indicates that the position updates are becoming denser and that the trend is continuous.
[0093] The preset threshold for the number of consecutive windows is the minimum requirement for confirming that the contraction trend is not accidental. Monotonically decreasing means that within the consecutive backtracking window segment where the contraction determination is true, the value of the subsequent window is never greater than the value of the previous window. The trend indicator of update density is the final output signal describing the systematic trend of the target object's position update pattern towards density.
[0094] For example, identify all consecutive backtracking window segments with a true contraction flag, and find the longest consecutive segment. Determine if the number of backtracking windows contained in this longest consecutive segment exceeds a preset threshold for the number of consecutive windows. If it does, extract the position update density value sequence of each backtracking window within this consecutive segment, and check if the sequence satisfies the condition that subsequent values are no greater than previous values, indicating a monotonically decreasing trend. If both the consecutive number and monotonically decreasing conditions are met, the module generates an indicator of the update density change trend and outputs it in conjunction with the current timestamp.
[0095] Step S350: Input the spatial distribution contraction trend indicator, the path detourness increase trend indicator, and the update density change direction indicator into the tendency fusion logic. The tendency fusion logic predefines the residence tendency level corresponding to different combinations of the three indicators.
[0096] In one implementation, the spatial distribution shrinkage trend indicator generated in step S320, the path detourness increase trend indicator generated in step S330, and the update density change direction indicator generated in step S340 are used as three Boolean input signals to the input of the tendency fusion logic. The tendency fusion logic is internally implemented as a decision tree or lookup table structure. For example, the root node of the decision tree first determines whether the update density change direction indicator is true. If it is true, it enters the left subtree to continue determining the path detourness increase trend indicator. If both the path detourness increase trend indicator and the spatial distribution shrinkage trend indicator are true, the highest dwell tendency level is output. If only the update density change direction indicator and the spatial distribution shrinkage trend indicator are true, and the path detourness increase trend indicator is false, the second highest dwell tendency level is output. And so on, each path from the root to the leaf in the decision tree corresponds to a combination of indicators, and the leaf node outputs the corresponding dwell tendency level.
[0097] Step S360: Output the residence tendency level that matches the current combination of the three indications from the tendency fusion logic, as a description of the residence tendency of the target object in the current area.
[0098] The dwell tendency description is the dwell tendency level finally output by the tendency fusion logic in step S350. This level can be directly used as a description of the dwell tendency of the target object in the current area for decision-making in step S400. The output dwell tendency description is an enumerated value with semantic labels or level numbers, clearly indicating whether the current target object is more inclined to continue roaming or is about to enter a stop-and-listen state.
[0099] Step S400: When the dwell tendency description reaches the preset tendency benchmark, and the output progress of the first audio content reaches any semantic breakpoint in the semantic segment division marker, the end of the output sampling sequence of the first audio content is continuously spliced with the beginning of the output sampling sequence of the second audio content in the output buffer to form an uninterrupted audio stream, and the directional sound source is controlled to output the second audio content.
[0100] In one implementation, step S400 may specifically include the following steps S410 to S460: Step S410: When the dwell tendency description reaches the tendency benchmark, a dwell stability indication is generated, and the output progress of the first audio content is continuously monitored to see if it touches any semantic breakpoint in the semantic segment division marker.
[0101] The dwell stability indicator is an internal event flag that is set when the dwell tendency description obtained from the event queue reaches the tendency benchmark, indicating that the target object has reached a dwell stability state, allowing and necessitating switching of audio content at appropriate semantic boundaries. Monitoring is the process by which the audio playback engine's sampling position polling thread continuously compares the current audio hardware pointer position with the temporal position of the semantic segment division marker point using sampling interrupts or timer interrupts.
[0102] For example, the system continuously reads the latest dwell tendency description record from the event queue. When the read dwell tendency level is greater than or equal to the tendency benchmark, the dwell stability indicator register is set. After detecting that the dwell stability indicator register is set, the playback control thread enables the semantic breakpoint listening flag and begins comparing the current output progress with the timestamps in the semantic segment division marker set in each sampling interrupt service routine. When the comparison finds that the time difference between the current output progress and a certain semantic breakpoint is less than the duration of one sampling period, it is determined that the semantic breakpoint has been reached, and a semantic breakpoint arrival event is generated.
[0103] Step S420: At the instant when the output progress reaches the semantic breakpoint, backtrack the sampling point of the preset splicing preprocessing length as the splicing preparation start point, and simultaneously acquire the trailing sampling sequence of the first audio content and the leading sampling sequence of the second audio content from the splicing preparation start point.
[0104] The splicing preprocessing length refers to the number of tail samples of the first audio content that need to be pre-trimmed before the semantic breakpoint, and the number of head samples of the second audio content that need to be pre-trimmed after the semantic breakpoint, providing a data time window for subsequent amplitude envelope analysis and crossfading. The trailing sample sequence is a segment of first audio content samples from the splicing preparation start point to the original end point or switching point of the first audio content. The preamble sample sequence is a segment of second audio content head samples of the same length or corresponding to the trailing sample sequence, starting from the output start position.
[0105] When a semantic breakpoint is reached, the read pointer of the current output buffer is immediately frozen. Using the current read pointer as a reference, a jump back is made to the number of sample points corresponding to the preset concatenation preprocessing length, and the position after the jump is recorded as the concatenation preparation start point. Subsequent sample points are continuously read from the circular buffer of the first audio content via direct memory access, starting from the concatenation preparation start point, until the current read pointer position, forming a trailing sample sequence. Simultaneously, starting from the beginning of the prefetched second audio content data block in the output buffer, sample points of the same length as the trailing sample sequence are read, forming a preamble sample sequence. Both sequences are stored in separate buffers.
[0106] Step S430: Extract the amplitude envelopes of the signals from the trailing sample sequence of the first audio content and the leading sample sequence of the second audio content, respectively, to obtain the trailing envelope contour and the leading envelope contour.
[0107] The amplitude envelope is a smooth curve reflecting the instantaneous amplitude change of an audio signal in the time domain. Stripping the amplitude envelope is typically achieved through full-wave rectification followed by low-pass filtering or by taking the modulus of the Hilbert transform. The tail envelope contour is the amplitude envelope of the tail sample sequence of the first audio content; its overall trend is typically a gradual decrease as the speech sentence naturally ends. The leading envelope contour is the amplitude envelope of the leading sample sequence of the second audio content; its overall trend is a gradual increase from silence or extremely low amplitude to normal volume.
[0108] In implementation, an envelope extraction unit can be used to extract the absolute value of each sample point in the trailing sample sequence. This absolute value sequence is then fed into a digital low-pass filter with an extremely low cutoff frequency. The filter employs a cascaded integrator-comb filter followed by a compensated finite impulse response filter to balance smoothing and group delay performance. The filtered output yields the trailing envelope contour. Simultaneously, the envelope extraction unit processes the leading sample sequence using the exact same procedure to obtain the leading envelope contour. The sampling rates of both envelope contours remain consistent with the original audio sampling rate.
[0109] Step S440: Extrapolate and extend the trailing envelope contour to generate a naturally decaying extension segment of the trailing envelope; perform initial prediction on the leading envelope contour to generate a predicted climbing segment of the leading envelope.
[0110] Extrapolation extends the trailing envelope by predicting its continued attenuation trajectory beyond the actual trailing data endpoint based on the attenuation slope and shape near the trailing end, thus avoiding auditory abruptness caused by premature muting. The natural attenuation extension segment of the trailing envelope is a virtual continuation of the trailing envelope on the time axis generated by extrapolation. Initial prediction infers the virtual amplitude change trajectory before the starting point based on the rise rate and initial amplitude characteristics of the leading envelope contour in the initial segment, resulting in a leading initial prediction curve that smoothly rises from zero amplitude. The expected rise segment of the leading envelope is a virtual warm-up segment of the leading envelope before the junction point, generated by initial prediction.
[0111] For example, the envelope extension unit takes a segment of sampling points at the tail of the trailing envelope contour and performs linear or exponential fitting. The exponential fitting method involves converting the tail data to the logarithmic domain, performing least-squares linear fitting in the logarithmic domain to obtain the decay slope, and then extending this decay slope forward along the time axis to generate an exponential decay curve, which serves as the natural decay extension segment of the trailing envelope. For the leading envelope contour, the envelope extension unit takes a portion of sampling points from its starting segment, similarly performs linear fitting in the logarithmic domain, and extends it in the negative direction of the time axis based on the climbing slope to generate an exponentially rising curve, thus obtaining the expected climbing segment of the leading envelope.
[0112] Step S450: In the overlapping area of the natural decay extension segment of the trailing envelope and the expected climbing segment of the leading envelope, a smooth mixing interval is defined with the intersection of the two envelope curves as the center, and amplitude cross-variable mixing is performed within the smooth mixing interval to obtain the envelope control curve of the mixing transition segment.
[0113] In one implementation, step S450 may specifically include the following steps S451 to S456: Step S451: In the overlapping region of the natural decay extension of the trailing envelope and the expected rise of the leading envelope, search along the time axis for the intersection of the two envelope curves, and take the intersection as the amplitude mixing center point.
[0114] The amplitude mixing center point is the symmetrical center moment of the crossover transition. Around this moment, the mixing weights have completed exactly half of the transfer, meaning the contribution of the trailing envelope and the leading envelope each account for half. Searching for the intersection of two envelope curves involves comparing the amplitude values of the two envelope curves point by point within the overlapping region to find the sampling position with the smallest absolute value of the amplitude difference, or locating the time point where the amplitudes are equal through linear interpolation.
[0115] For example, starting from the beginning of the overlapping region, the absolute value of the amplitude difference between the naturally decaying extension of the trailing envelope and the expected rise of the leading envelope at the current moment is calculated by stepping through the sampling points. The sampling point where the absolute value of the amplitude difference reaches its minimum during the traversal is recorded. If there are two adjacent sampling points with inverted amplitude difference signs, the precise intersection point where the amplitude difference is zero is calculated between these two sampling points by linear interpolation, and this intersection point is determined as the amplitude mixing center point.
[0116] Step S452: Using the amplitude mixing center point as a reference, extend the preset mixing half-interval duration forward and backward respectively, and take the closed interval from the end point of the forward extension to the end point of the backward extension as the smooth mixing interval.
[0117] The half-interval duration is a preset time length parameter that determines half of the total duration of the crossover transition. It determines the smoothness of the transition; the longer the half-interval duration, the smoother the transition. The starting point of the smooth blending interval is determined by the endpoint reached by extending the half-interval duration forward from the amplitude blending center point, and the ending point is determined by extending the half-interval duration backward from the amplitude blending center point.
[0118] Optionally, a preset half-interval duration parameter can be read. Using the time coordinate of the amplitude mixing center point as a reference, the half-interval duration is subtracted to obtain the start time of the smooth mixing interval; using the time coordinate of the amplitude mixing center point as a reference, the half-interval duration is added to obtain the end time of the smooth mixing interval. The closed interval between the start time and the end time is defined as the smooth mixing interval.
[0119] Step S453: Within the smooth mixing interval, compare the envelope amplitude values of the naturally decaying extension segment of the trailing envelope with the expected climbing segment of the leading envelope at each time step to identify the amplitude dominance relationship between the two curves at each time step.
[0120] The amplitude dominance relationship refers to whether the amplitude of the trailing envelope curve or the amplitude of the leading envelope curve is larger at any time within the smooth mixing interval. This relationship is used to guide whether additional gain compensation is needed for one side during subsequent weighted mixing to avoid a dip in the total loudness.
[0121] Within the smooth mixing interval, the amplitude values of the tail envelope's naturally decaying extension and the leading envelope's expected ascent are read point by point. At each sampling point, the two amplitude values are compared. If the tail envelope amplitude is greater than the leading envelope amplitude, the current time is marked as tail-dominated; if the leading envelope amplitude is greater than the tail envelope amplitude, it is marked as leading-dominated; if they are equal, it is marked as balanced.
[0122] Step S454: At the beginning of the smooth mixing interval, the mixing weight is fully allocated to the naturally decaying extension of the trailing envelope; as time progresses towards the amplitude mixing center point, the mixing weight ratio of the expected climbing segment of the leading envelope is gradually increased; at the amplitude mixing center point, the mixing weights of the two envelopes are equal; after passing the amplitude mixing center point, the mixing weight ratio of the expected climbing segment of the leading envelope continues to increase until the mixing weight at the end of the smooth mixing interval is fully allocated to the expected climbing segment of the leading envelope.
[0123] In one implementation, step S454 may specifically include the following steps S4541 to S4546: Step S4541: Obtain the start time, amplitude mixing center point time, and end time of the smooth mixing interval. Define the period from the start time to the amplitude mixing center point time as the mixing front segment and the period from the amplitude mixing center point time to the end time as the mixing back segment.
[0124] The first phase of the mixing process is a gradual transition from the tail envelope completely dominating to an equal-weighted mixing of the two envelopes. The second phase is a gradual transition from an equal-weighted mixing of the two envelopes to a completely dominant leading envelope. The two phases smoothly connect at the amplitude mixing center point. For example, read the start time, amplitude mixing center point time, and end time of the smooth mixing interval. Mark the time period from the start time to the amplitude mixing center point time as the first phase of the mixing process, and mark the time period from the amplitude mixing center point time to the end time as the second phase of the mixing process. The duration of both time periods is equal to the duration of the half-interval of the mixing process.
[0125] Step S4542: In the mixing front end, determine the mixing progress position at the current moment based on the proportion of the time distance between the current moment and the starting moment to the total duration of the mixing front end. Gradually decrease the mixing weight of the tail envelope natural decay extension segment from the full weight at the starting end, while gradually increase the mixing weight of the leading envelope expected climbing segment from the zero weight at the starting end. The increase or decrease magnitude corresponds linearly with the mixing progress position.
[0126] The blending progress position is a proportional value between 0 and 1, representing the percentage of total time already progressed in the initial blending phase. Linear correspondence means that the change in blending weights is linearly related to the blending progress position; that is, the decrease in the tail envelope weight and the increase in the leading envelope weight are both linear functions of time.
[0127] In the specific execution of step S4542, weight calculation is performed point-by-point within the mixing front segment. For the current sampling point, the time difference between the current time and the start time of the mixing front segment is calculated, and then divided by the total duration of the mixing front segment to obtain the mixing progress position. The mixing weight of the naturally decaying extension segment of the trailing envelope is set to the value of 1 minus the value of the mixing progress position, and the mixing weight of the expected climbing segment of the leading envelope is set to the value of the mixing progress position. The sum of the two is always 1.
[0128] Step S4543: At the amplitude mixing center point, set the mixing weight of the naturally decaying extension of the trailing envelope and the mixing weight of the expected climbing segment of the leading envelope to the same value.
[0129] The amplitude mixing center point corresponds to a mixing progress position of 0.5. At this point, the mixing weights of the trailing envelope and the leading envelope are both precisely 0.5, and the envelope contributions of the two signals are completely equal.
[0130] During the specific execution of step S4543, it is detected that the current sampling time has reached the amplitude mixing center point. The mixing weight of the natural decay extension segment of the trailing envelope is directly set to a value of 0.5, and the mixing weight of the expected climbing segment of the leading envelope is also set to a value of 0.5.
[0131] Step S4544: In the later stage of mixing, the mixing progress position at the current moment is determined based on the proportion of the time distance between the current moment and the moment of the amplitude mixing center point to the total duration of the later stage of mixing. The mixing weight of the natural decay extension segment of the trailing envelope is gradually reduced from the same value, while the mixing weight of the expected climbing segment of the leading envelope is gradually increased from the same value. The increase or decrease magnitude is linearly correlated with the mixing progress position.
[0132] The calculation method for the mixing progress position in the later stage of mixing is the same as that in the earlier stage of mixing, except that the starting reference point is changed to the amplitude mixing center point, the weight of the trailing envelope continues to decrease to zero, and the weight of the leading envelope continues to increase to full weight.
[0133] For example, the time difference between the current moment and the amplitude mixing center point is calculated for each sampling point in the mixing phase, and divided by the total duration of the mixing phase to obtain the mixing progress position. The mixing weight of the natural decay extension of the trailing envelope is set to 0.5 multiplied by 1 minus the difference in mixing progress position, and the mixing weight of the expected rise of the leading envelope is set to 0.5 plus 0.5 multiplied by the mixing progress position, ensuring that the weight of the trailing envelope reaches zero and the weight of the leading envelope reaches full weight at the end of the mixing phase.
[0134] Step S4545: During the advancement of the mixing front and mixing back segments, in real time, detect whether the envelope amplitude value of the naturally decaying extension segment of the trailing envelope has decayed to below the inaudible limit at the current moment. If so, immediately set the mixing weight of the expected climbing segment of the leading envelope to full weight and end the mixing transition in advance.
[0135] The inaudible threshold is an extremely small envelope amplitude value set based on the human hearing threshold or the system noise level. When the trailing envelope decays below this threshold, continuing to perform crossfade is no longer audible, and the system can directly switch to the prelead envelope. Ending the mixing transition early is to avoid applying incomplete weights to the prelead signal and weakening its loudness after the trailing has completely faded.
[0136] During the specific execution of step S4545, while calculating the weight each time, the envelope amplitude value of the naturally decaying extension segment of the trailing envelope at the current moment is read. If the amplitude value is already less than the inaudible limit, the gradation process is immediately terminated, the mixing weight of the naturally decaying extension segment of the trailing envelope at the current moment and all remaining moments in the smooth mixing interval is forcibly set to zero, the mixing weight of the expected climbing segment of the leading envelope is forcibly set to full weight, and the mixing transition is marked as ending early.
[0137] Step S4546: During the advancement of the mixing front and mixing back segments, it is detected in real time whether the envelope amplitude value of the expected climbing segment of the leading envelope has not yet climbed above the audible limit at the current moment. If so, the mixing weight of the naturally decaying extension segment of the trailing envelope continues to dominate until the envelope amplitude value of the expected climbing segment of the leading envelope crosses the audible limit.
[0138] The audible limit is an envelope amplitude threshold set according to the human hearing threshold. When the amplitude of the leading envelope is below this limit, even if it is given a high weight, it cannot produce effective auditory perception. Maintaining the dominance of the trailing envelope can avoid the occurrence of brief auditory voids. Crossing the audible limit refers to the moment when the amplitude of the leading envelope increases to be greater than the audible limit.
[0139] When calculating weights, the envelope amplitude of the expected climbing segment of the leading envelope at the current moment is monitored simultaneously. If this amplitude has not yet reached the audible limit, the current mixing progress is temporarily frozen, and the weight ratio of the expected climbing segment of the leading envelope is not further increased, maintaining the high proportion of the naturally decaying extension segment of the trailing envelope at the current moment. The freeze is lifted and the normal linear gradual change process of the weights is resumed once the envelope amplitude of the expected climbing segment of the leading envelope is detected to be greater than the audible limit.
[0140] Step S455: At each moment, the envelope amplitude values of the naturally decaying extension segment of the trailing envelope and the expected climbing segment of the leading envelope are weighted and mixed according to the mixing weight ratio to generate the mixed envelope amplitude value at that moment.
[0141] Weighted mixing refers to multiplying the tail envelope weight by the current amplitude value of the tail envelope, adding the preceding envelope weight multiplied by the current amplitude value of the preceding envelope, and calculating the weighted sum, which is the mixed envelope amplitude value. The mixed envelope amplitude value is the target envelope amplitude obtained after weighted fusion at each sampling time within the smooth mixing interval.
[0142] Within the smooth mixing interval, for each sampling point, the tail envelope weight determined in step S454 is multiplied by the envelope amplitude value of the naturally decaying extension segment of the tail envelope at that time, and the leading envelope weight is multiplied by the envelope amplitude value of the expected climbing segment of the leading envelope at that time. The two products are then added together to obtain the mixed envelope amplitude value at that time.
[0143] Step S456: Collect the mixing envelope amplitude values at all times within the smooth mixing interval to form the envelope control curve of the mixing transition segment.
[0144] The envelope control curve of the mixing transition section is a control sequence defined within the smooth mixing interval time range, specifying the instantaneous amplitude envelope target value that the mixed output audio should have at each time point. This curve will be directly used to modulate the amplitude of the original audio sampling points.
[0145] For example, all the mixed envelope amplitude values generated in the order of sampling points within the smooth mixing interval are sequentially written into an array with the same length as the number of sampling points in the smooth mixing interval. This array is the envelope control curve of the mixing transition segment.
[0146] Step S460: Apply the envelope control curve of the mixed transition segment to the corresponding sampling points of the first audio content trailing sampling sequence and the second audio content leading sampling sequence to synthesize the transition segment. Then, sequentially concatenate the sampling sequence of the first audio content before the start of the transition preprocessing, the transition segment, and the sampling sequence of the second audio content after the transition interval to obtain a seamless audio stream.
[0147] The synthesized transition segment is formed by multiplying the corresponding sampling points of the first audio content trailing sample sequence and the second audio content preamble sample sequence in the time domain by the actual tail gain and the actual preamble gain calculated from the blending transition segment envelope control curve, and then summing them to form a short transition audio sample data segment. Sequential concatenation refers to joining the three sample data segments end to end in chronological order and writing them into the output buffer to form a continuous data stream.
[0148] First, based on the relationship between the envelope control curve of the mixed transition segment and the original two envelope curves, the tail gain curve and the leader gain curve are decomposed. The tail gain curve is equal to the envelope control curve divided by the amplitude value of the tail envelope profile at the corresponding time, and the leader gain curve is equal to the envelope control curve divided by the amplitude value of the leader envelope profile at the corresponding time. Each sampling point of the tail sampling sequence is multiplied by the tail gain curve value at the corresponding time, and each sampling point of the leader sampling sequence is multiplied by the leader gain curve value at the corresponding time. The products are then added together to obtain the audio sampling data of the transition segment. The portion of the first audio content output buffer before the splicing preparation point, the synthesized transition segment, and the portion of the second audio content after the smooth mixing interval are concatenated in chronological order and written into the output circular buffer of the audio playback front end to form a continuous, uninterrupted audio stream, which is then submitted to the directional sound source output.
[0149] Step S500: Use the second audio content as the new first audio content, retain the audio data that has not yet been output in the output buffer as the current playback data source, reset the prefetch state of the output buffer to idle, and return to the step of parsing the semantic segment division markers pre-embedded in the first audio content during the output process of the first audio content, until the sensor stops generating the sensing signal of the target object.
[0150] In one implementation, step S500 may specifically include the following steps S510-S560: Step S510: After completing the output switching, overwrite the audio identifier of the second audio content with the audio identifier of the current first audio content, and replace the semantic segmentation markers carried in the second audio content with the currently available semantic segmentation markers.
[0151] The audio identifier is an internal code used in the system to uniquely identify the currently playing audio content resource. The overwrite operation refers to updating the field recording the identifier of the first audio content in the current playback context structure in memory to the identifier of the second audio content, thus transferring the playback identity. Replacing semantic segmentation markers means updating the list of semantic segmentation markers used by the current playback monitoring module to a list of semantic breakpoints parsed from the second audio content, allowing subsequent progress monitoring to be based on the semantic structure of the new content.
[0152] After confirming that the uninterrupted audio stream has successfully started output, a role switch operation is performed. The manager accesses the current playback session control block and rewrites the value of the current audio identifier field to the file index number of the second audio content in the preset audio content library. At the same time, the manager reads the pre-embedded semantic segmentation marker data from the metadata header of the second audio content, parses it into an ordered list of timestamps, and uses this list to overwrite the semantic segmentation marker array stored in the playback session control block, completing the semantic breakpoint update.
[0153] Step S520: Retain the audio data that has not yet been output in the output buffer and continue to use it as the output data of the current first audio content. Clear the prefetch occupancy flag of the output buffer, adjust the write pointer to the starting address of the free prefetch area, and keep the read pointer unchanged at the current playback position.
[0154] The output buffer prefetch occupancy flag is a semaphore or mutex flag used to prevent data contention between the prefetch thread and the playback thread. Clearing this flag releases the prefetch operation's exclusive right to write to the output buffer's write area. Adjusting the write pointer to the starting address of the free prefetch area allows the next round of prefetch operations to know where to begin writing new prefetched audio data; this address is typically immediately after the end of the second prefetched audio content block.
[0155] In the callback function after the switch is complete, no jump is made to the read pointer; it remains pointing to the current output position. The prefetch occupancy flag in the output buffer control structure is cleared. The last physical address occupied by the prefetched data of the second audio content in the output buffer is calculated, and this address plus an offset is used as the new write pointer. This write pointer points to the starting address of the free prefetch region, thus preparing the write position for the next prefetch operation.
[0156] Step S530: Continuously monitor the signal line status of the sensor. When the signal line status continuously drops from an active level to an inactive level and remains inactive for a duration exceeding a preset loss confirmation period, generate a sensor interruption event.
[0157] Signal line status refers to the logic level on the digital input / output pins connecting the sensor and the audio processing system. A valid level can be high or low, depending on the interface protocol, indicating that the sensor is continuously tracking the target object. An invalid level is the default signal line status after the target object leaves the sensing range. The preset loss confirmation period is to filter out momentary signal loss caused by brief obstruction of the target object or instantaneous sampling jitter of the sensor; only when the signal level continuously drops for more than this period is the target object truly confirmed lost. A sensing interruption event is an internal interrupt message triggered when the target object disappears from the sensing system's field of view.
[0158] Step S540: During the duration of the sensing interruption event, extract the last recorded position description and the last generated dwell tendency description before the interruption, and generate a set of inferred positions based on the inertial continuation trajectory along the displacement direction before the interruption according to the position description.
[0159] In one implementation, step S540 may include the following steps S541 to S546: Step S541: Extract the position description sequence of the last segment recorded before the interruption, calculate the displacement vector between adjacent position descriptions in the last segment, and perform direction concentration analysis on the directions of all displacement vectors in the last segment to determine the dominant displacement direction.
[0160] The final segment position description sequence is a subset of the position descriptions within the last time window before the sensing interruption. Displacement vector direction concentration analysis measures whether multiple displacement vectors within this final segment point in similar directions. If the displacement vectors are concentrated in a specific distribution, it indicates that the target object had a clear direction of movement before the interruption, and the main displacement direction can be extracted.
[0161] Specifically, the most recently written position description is extracted from the time-sorted set of position descriptions as the last segment. The displacement vector formed by every two adjacent position descriptions within this segment is calculated, yielding several displacement direction angles. The entire circumference angle is divided into several sectors, and a direction histogram is constructed. The frequency of displacement vectors within each sector is statistically analyzed. The center angle of the sector with the highest frequency is taken as the dominant displacement direction. If the frequency proportion of the sector with the highest frequency in the histogram exceeds a preset dominant proportion, the dominant displacement direction is confirmed to be valid; if the frequency distribution is too scattered and no sector exceeds the preset dominant proportion, it is determined that there is no dominant direction. In this case, the dominant displacement direction can be set as the direction of the last displacement vector.
[0162] Step S542: Extract the length sequence of displacement vectors in the last segment, calculate the decay or growth trend of displacement vector length, and obtain the evolution trend of displacement length.
[0163] The displacement vector length sequence is an ordered sequence of the lengths of each displacement vector within the final segment. The trend of displacement length evolution is a description of the trend obtained by fitting the direction of change of this length sequence. For example, a gradually shortening length indicates that the target object is decelerating, while a gradually increasing length indicates that the target object is accelerating.
[0164] For example, extract the length values of each displacement vector within the final segment and arrange them in chronological order. Perform a linear fit on this length sequence and calculate the slope of the fitted line. If the slope is greater than 0, the displacement length evolution trend is increasing; if the slope is less than 0, the displacement length evolution trend is decreasing; if the slope is approximately equal to 0, the displacement length evolution trend is flat.
[0165] Step S543: Using the last recorded position before the interruption as the starting point of inertial continuation, the main displacement direction as the continuation direction, and the evolution of displacement length as the basis for step size change, the inference is generated moment by moment along the time progression direction, and the predicted position at each inference moment is generated.
[0166] The inertial continuation starting point is the last acquired two-dimensional or three-dimensional spatial coordinates of the target object before the interruption. Time-by-time simulation along the time progression direction refers to simulating the recursive process of the target object's continued movement from the inertial continuation starting point, according to a preset time step. During the simulation, the step size of each time step is determined by the evolution of the displacement length, and the step size can be gradually increasing or decreasing.
[0167] For example, the time step for the simulation is set to be equal to the average sampling period of the position description acquisition. The coordinates of the inertial continuation starting point are used as the initial coordinates for the simulation. The fitted linear equation of the displacement length evolution trend is extracted, the average step size of the first step of the simulation is calculated, and the displacement of this step size is superimposed on the initial coordinates along the main displacement direction to obtain the predicted position at the first simulation moment. Subsequently, the step size of the next step is adjusted according to the evolution trend, and the position superposition process is repeated to gradually generate the predicted positions at each simulation moment.
[0168] Step S544: Continue the simulation until the simulation time exceeds the duration of the induced interruption event. Collect all the predicted positions generated in this process in chronological order to form a set of predicted positions.
[0169] The simulation terminates when the cumulative time span of the simulation is equal to or exceeds the duration of the induced interruption event from its occurrence to the present. The predicted location set contains all possible spatiotemporal trajectory sampling points of the target object from the start of the predicted interruption to the current moment.
[0170] After each new predicted location is generated, the simulation time step is accumulated, and it is checked whether the accumulated simulation time is greater than or equal to the duration of the interrupted event. If not, the simulation continues to the next one; if it has been reached or exceeded, the simulation stops, and all generated predicted locations are sorted into a predicted location set in ascending order of simulation time.
[0171] Step S545: Perform an inclusion judgment between each predicted position in the predicted position set and the spatial boundary of the current region, count the first time when consecutive predicted positions exceed the spatial boundary, and use the difference between the first time and the interruption start time as the disengagement duration.
[0172] The spatial boundary is the predefined geometric boundary of the current region, which can be a rectangular bounding box, a polygonal boundary, or a closed polygon composed of multiple line segments. Inclusion determination is a geometric calculation to determine whether a point is inside the polygon boundary, using either the ray method or the cross product method. Escape duration is the time elapsed from the occurrence of a sensing interruption to the inferred target object first crossing the region boundary. If all inferred positions have not exceeded the boundary, the escape duration is a value greater than the already sustained duration.
[0173] Specifically, each predicted location in the predicted location set is traversed. For each predicted location, an inclusion test is performed within the polygon, and the cross product direction of its coordinates is checked sequentially with the coordinates of the boundary vertices of the current region. When the inclusion test result for a predicted location is first detected as false, the prediction timestamp corresponding to that predicted location is recorded. The difference between this prediction timestamp and the interruption start time is calculated as the detachment duration. If no boundary exceedance is found after traversing all predicted locations, the detachment duration is set to a maximum value greater than the duration of the detected interruption event.
[0174] Step S546: When all predicted positions in the predicted position set do not exceed the spatial boundary of the current area, the separation duration is set to a value greater than the duration of the sensing interruption event.
[0175] This step ensures that if the inertial trajectory of the target object remains completely within the current area, it is determined that the object is still residing within the current area and will not trigger the acoustic weakening process.
[0176] After completing the inclusion judgment traversal, if the out-of-bounds event has never occurred, the value of the departure duration variable is set to the duration of the sensing interruption event plus a preset large time increment to ensure that the departure duration is always greater than the duration in subsequent comparisons.
[0177] Step S550: Overlay and compare the inferred location set with the spatial boundary of the current area to determine the time required for the inferred location to leave the boundary of the current area. When the time required to leave the boundary exceeds the duration of the sensing interruption event, determine that the target object is still residing in the current area and maintain the normal output of the current first audio content and semantic breakpoint monitoring.
[0178] The overlay comparison refers to the inclusion judgment and separation time extraction operations performed in steps S540 to S545. Determining that the target object is still within the current area means that based on the condition that the inertial continuation trajectory has not left the area boundary, it is inferred that the target object is likely still within the area, and the system maintains normal operation mode accordingly.
[0179] The system compares the separation duration with the duration of the sensor interruption event. If the separation duration is greater than the duration, the target object is considered to still be within the current spatial range. No departure confirmation is generated, the current playback session remains unchanged, and the semantic breakpoint monitoring and prefetching switching logic for the second audio content continues. Simultaneously, it waits for the sensor signal to recover. If the sensor signal recovers within a short time, the inferred location set is immediately discarded, and the actual location is resynchronized.
[0180] Step S560: When the disconnection time is not greater than the duration of the sensing interruption event, record the most recent semantic breakpoint of the current first audio content that has not yet been output, start the output sound energy weakening process at the semantic breakpoint until complete silence, clear the position description time sorting set, and return to the initial state of the sensing signal generated by the receiving sensor.
[0181] In one implementation, step S560 may specifically include the following steps S561 to S566: Step S561: When the disconnection duration is not greater than the duration of the sensing interruption event, in the currently available semantic segment division markers, search along the time axis for the first semantic breakpoint whose time position is later than the current output progress, and use this semantic breakpoint as the weakening start point.
[0182] Searching in ascending timeline refers to retrieving the list of semantic segment markers in ascending order of time, finding the first marker whose timestamp is greater than the current audio playback hardware pointer time. The decay start point is the time when the acoustic decay processing begins to take effect; selecting the nearest future semantic breakpoint ensures that acoustic decay also occurs at a semantically complete and natural pause. For example, the current output progress time in the current playback session control block is read and compared with each timestamp in the semantic segment marker array. Comparisons are made sequentially in ascending array index; when the first timestamp is found to be greater than the current output progress time, the search stops, and this timestamp is recorded as the decay start point.
[0183] Step S562: Starting from the weakening start point, start the sound energy weakening process for a preset weakening span. The sound energy weakening process causes the output sound energy envelope to continuously decay from the current amplitude level to the silent level within the preset weakening span.
[0184] The preset decay span is the time required for the sound energy to attenuate from the normal playback level to silence. The sound energy decay processing is achieved by decreasing the digital gain coefficient point by point, and the gain coefficient changes with time according to the preset decay curve.
[0185] The fading process is triggered when the output progress reaches the fading start point. The current value of the digital volume gain register is used as the initial gain. According to the sampling cycle, the gain decrease step size for each sampling point is calculated based on the preset fading span and the initial gain. The step size is calculated using a linear decrease method, meaning that the gain decreases by a fixed, small increment for each sampling point. Each sampling point is multiplied by the current instantaneous gain value before being output, causing the sound energy envelope to decay linearly until the gain returns to zero, reaching a silent level.
[0186] Step S563: During the execution of the acoustic energy fading process, continue to monitor the status of the sensor's signal line. When the sensor signal recovers within the preset fading span and the first position description collected after the signal recovery falls into the current area, interrupt the acoustic energy fading process.
[0187] Interrupted audio fading processing refers to the process where, during the fading process, the system decides to abandon the silence and resume normal audio output because the induction signal recaptures the target object and confirms that it is still within the area.
[0188] The signal monitoring thread continuously listens to the sensor signal line during the fading process. Once the signal line state is detected to recover from an invalid level to an active level within the preset fading span, and the first location description acquired and parsed after the sensing system recovers is confirmed by inclusion determination to be within the spatial boundary of the current region, the thread immediately sends a fading interruption request to the audio playback engine. The audio playback engine responds to the request, stops the gain reduction operation, and freezes the current gain value as the current remaining acoustic energy amplitude at the time of the fading interruption.
[0189] Step S564: After the interruption of the acoustic energy fading process, record the current remaining acoustic energy amplitude when the fading process is interrupted, and starting from the current remaining acoustic energy amplitude, continuously increase the output acoustic energy envelope from the current remaining acoustic energy amplitude to the normal output amplitude within the preset acoustic energy recovery span.
[0190] The preset gain recovery span is the time required for the gain to recover from its lower level at the point of gradual decline to the normal playback level. Continuous boost refers to the gain coefficient increasing linearly at a fixed slope until it recovers to unity gain. For example, the frozen current remaining sound energy amplitude value is read and used as the starting gain for the recovery. Based on the preset gain recovery span and the target full gain value, the gain increment step size for each sampling point is calculated. Before each subsequent sampling point output, the gain value is increased by one step, and the sample point is multiplied by the incremented gain before being sent to the digital-to-analog converter until the gain reaches the full gain value, and the sound energy is fully restored to the normal output amplitude.
[0191] Step S565: After the sound energy recovery is completed, the total duration of the sound energy weakening and the sound energy recovery during the induction interruption is used as the synchronization compensation amount to perform a rollback compensation on the output progress of the current first audio content, so that the output progress is re-aligned with the real timeline.
[0192] Synchronization compensation is used because during the fading and recovery phases, although the audio output is not interrupted, the playback progress continues. The target audience may miss some audio content, requiring the playback pointer to be adjusted backward to compensate for the missed content, or conversely, to jump forward to maintain the correlation with real time. Waiting back compensation refers to subtracting the synchronization compensation amount from the current output progress time value of the audio playback, causing the playback to return to an earlier time position.
[0193] Specifically, the sum of the actual duration of the sound energy fading and the actual duration of the sound energy recovery is calculated as the synchronization compensation amount. The synchronization compensation amount is subtracted from the timestamp corresponding to the current output progress to obtain the compensated output progress time. The audio hardware playback pointer is then moved to the sampling point offset position corresponding to this compensated output progress time, so that the subsequent audio output continues from the compensated time point, thereby realigning with the auditory perception timeline of the target object.
[0194] Step S566: Using the aligned output progress and the recovered location description as a new starting point, restart the accumulation of the location description time sorting set and the generation of the dwell tendency description, and continue to execute the semantic breakpoint monitoring and second audio content switching loop.
[0195] Restarting the accumulation of the location description time sorting set means clearing the original location description time sorting set and starting from the first location description collected after recovery to begin accumulating location description data again. Continuing to execute the semantic breakpoint monitoring and second audio content switching loop means that the system fully recovers to the normal adaptive audio service loop logic with the target object present.
[0196] For example, clear all historical data from the location description time-ordered set, and use the first location description and its timestamp acquired after recovery as the first element of the set. The generation process of the dwell tendency description starts from scratch and re-executes the gradual accumulation of the backtracking window. At the same time, continue to monitor the semantic segment division markers, and wait for the dwell tendency description to reach the tendency benchmark again. Then, at the next semantic breakpoint, prefetch and switch to the next version of audio content. The entire system enters a new round of cyclic service from steps S200 to S500 until the sensing signal completely disappears again and it is confirmed that the target object has left the area boundary.
[0197] This invention also provides a server, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the steps in the regional sound field adjustment method based on an inductive directional sound source provided in this invention.
[0198] Please see details. Figure 4 This is a schematic diagram of the structure of a server provided in an embodiment of the present invention. Figure 4 As shown, the server 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the server 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. Figure 4 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.
[0199] exist Figure 4 In the server 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the methods provided in the above embodiments.
[0200] It should be understood that the server 1000 described in this embodiment of the invention can execute the foregoing text. Figure 2 The implementation principle and beneficial effects of the regional sound field adjustment method based on inductive directional sound sources described in the corresponding embodiments will not be repeated here.
Claims
1. A method for regional sound field adjustment based on an inductive directional sound source, characterized in that, include: Receive a sensing signal generated by a sensor corresponding to a target object and a location description of the area where the target object is located; retrieve a first audio content associated with the location description based on the sensing signal and the location description; and control a directional sound source to output the first audio content to the area where the target object is located. During the output of the first audio content, the pre-embedded semantic segmentation markers within the first audio content are parsed to obtain the current output progress of the first audio content, and the remaining output duration of the nearest semantic breakpoint is determined based on the semantic segmentation markers and the current output progress. Using the remaining output duration and the position description as input, determine the second audio content in the preset audio content library, and prefetch the second audio content into the output buffer; The location description sequence of the target object generated by the sensor is continuously collected to form a time-ordered set of location descriptions; the location fluctuation analysis is performed on the time-ordered set of location descriptions to generate a description of the dwelling tendency of the target object in the current area, and the dwelling tendency description characterizes the strength of the tendency of the target object to change from a wandering state to a stationary state. When the dwell tendency description reaches the preset tendency benchmark, and the output progress of the first audio content reaches any semantic breakpoint in the semantic segment division marker, the end of the output sampling sequence of the first audio content is continuously spliced with the beginning of the output sampling sequence of the second audio content in the output buffer to form an uninterrupted audio stream, and the directional sound source is controlled to output the second audio content. The second audio content is used as the new first audio content. The audio data that has not yet been output in the output buffer is retained as the current playback data source. The prefetch state of the output buffer is reset to idle. The step of parsing the semantic segmentation markers pre-embedded in the first audio content during the output process of the first audio content is returned until the sensor stops generating the sensing signal of the target object.
2. The regional sound field adjustment method based on an inductive directional sound source according to claim 1, characterized in that, During the output process of the first audio content, the semantic segmentation markers pre-embedded in the first audio content are parsed to obtain the current output progress of the first audio content, and the remaining output duration of the nearest semantic breakpoint is determined based on the semantic segmentation markers and the current output progress. Using the remaining output duration and the position description as input, the second audio content is determined from a preset audio content library, and the second audio content is prefetched into the output buffer, including: The first audio content is segmented in the time domain to obtain several overlapping short-time analysis panes. The peak trend of sound pressure level and the change pattern of zero-crossing rate are extracted in each short-time analysis pane to form a zero-crossing rate-peak combination description sequence. In the zero-crossing rate-peak value combination description sequence, the silent interval segment where the zero-crossing rate is continuously lower than the silent zero-crossing threshold and the sound pressure level peak value is continuously lower than the silent energy threshold is locked, and the midpoint time of each silent interval segment is extracted to generate a set of silent segmentation points. Perform coherence analysis of harmonic components on the first audio content, detect the abrupt change position of the harmonic structure between adjacent short-time analysis panes, extract the time of each harmonic structure abrupt change, and generate a set of harmonic inflection points. The set of silent segmentation points and the set of harmonic inflection points are combined and merged along the time axis. When the time interval between the silent segmentation point and the harmonic inflection point is less than the preset fusion tolerance interval, one of them is removed; otherwise, both are retained to form a coarsely selected set of breakpoints. The coarsely selected set of breakpoints is filtered for contextual semantic coherence, and isolated breakpoints that appear in the middle of a coherent speech segment are removed to obtain the semantic segment division markers. Based on the remaining output duration and the location description, multiple versions of narrative audio corresponding to the location description are queried in the preset audio content library. According to the narrative branching link clues recorded in the first audio content, one version is selected from the multiple versions of narrative audio as the second audio content and stored in the output cache.
3. The regional sound field adjustment method based on an inductive directional sound source according to claim 2, characterized in that, The step of performing coherence analysis of harmonic components on the first audio content, detecting abrupt changes in harmonic structure between adjacent short-time analysis panes, extracting the time of each harmonic structure abrupt change, and generating a set of harmonic inflection points includes: Multi-resolution spectral decomposition is performed on each short-time analysis pane of the first audio content to generate a spectral linear array composed of several frequency components, and the candidate harmonic frequency position set is locked in the spectral linear array based on the fundamental frequency prediction trajectory. Frequency component pairing is performed on the spectral linear array of adjacent short-time analysis panes. A one-to-one correspondence is established between the position of each candidate harmonic frequency in the preceding pane and the component with the smallest frequency offset in the following pane, forming a cross-window harmonic pairing link. The amplitude evolution curve and frequency offset curve of each harmonic component are traced one by one along the cross-window harmonic pairing link. The moment when the slope reverses in the amplitude evolution curve and the moment when the direction reverses in the frequency offset curve are detected are marked as amplitude anomaly points and frequency anomaly points, respectively. In the same cross-window harmonic pairing link, when the time interval between the amplitude anomaly point and the frequency anomaly point is less than the preset synchronization anomaly judgment interval, the time position is taken as the instability position of the link; otherwise, the independent instability positions corresponding to the amplitude anomaly point and the frequency anomaly point are retained respectively. The number of links with unstable positions in each cross-window harmonic pairing link at the boundary of the same short-time analysis pane is counted. When the ratio of the number of such links to the total number of cross-window harmonic pairing links exceeds the preset instability ratio threshold, it is determined that a harmonic structure abrupt change has occurred at the boundary of the current short-time analysis pane. The set of harmonic inflection points is generated by collecting all short-time analysis pane boundary moments that are determined to have caused harmonic structural abrupt changes, arranging them in chronological order, and removing consecutively repeating adjacent boundary moments.
4. The regional sound field adjustment method based on an inductive directional sound source according to claim 3, characterized in that, In the same cross-window harmonic pairing link, when the time interval between the amplitude anomaly point and the frequency anomaly point is less than the preset synchronization anomaly judgment interval, this time position is taken as the instability position of the link; otherwise, the independent instability positions corresponding to the amplitude anomaly point and the frequency anomaly point are retained respectively, including: For each cross-window harmonic pairing link, its amplitude evolution curve is traversed along the time axis. When an inflection point is detected where the slope of the amplitude evolution curve changes from positive to negative or from negative to positive, the time corresponding to that inflection point is recorded as a candidate point for amplitude anomaly. The candidate points of amplitude anomaly are screened, and the steepness of the slope change of the amplitude evolution curve before and after the point is calculated. Only the candidate points of amplitude anomaly whose slope change exceeds the preset steepness threshold are retained as the amplitude anomaly points. Traverse the frequency offset curve along the time axis of the same cross-window harmonic pairing link. When the direction of the frequency offset curve changes from shifting to high frequency to shifting to low frequency or from shifting to shifting to high frequency, record the time corresponding to the direction switching point as a candidate point for frequency anomaly. The candidate frequency anomalies are confirmed, and the change in the offset rate of the frequency offset curve before and after the point is calculated. Only the candidate frequency anomalies whose change in offset rate exceeds the preset rate change threshold are retained as the frequency anomalies. The amplitude anomaly points and frequency anomaly points in the same cross-window harmonic pairing link are arranged in chronological order. When the time span between adjacent amplitude anomaly points and frequency anomaly points is less than the preset synchronization anomaly judgment interval, the midpoint of their time is taken as the instability position. When there are no frequency anomalies before or after the amplitude anomaly point with a time span smaller than the preset synchronization anomaly judgment interval, the amplitude anomaly point is taken as the independent instability location; when there are no amplitude anomalies before or after the frequency anomaly point with a time span smaller than the preset synchronization anomaly judgment interval, the frequency anomaly point is taken as the independent instability location.
5. The regional sound field adjustment method based on an inductive directional sound source according to claim 1, characterized in that, The continuous acquisition of the position description sequence of the target object generated by the sensor forms a time-ordered set of position descriptions with a temporal sequence. Perform location fluctuation analysis on the time-ordered set of location descriptions to generate a description of the target object's tendency to reside in the current area, including: The location description time sorting set is divided into multiple backtracking windows in chronological order. Each backtracking window contains several continuously collected location descriptions. Adjacent backtracking windows maintain a sliding progressive relationship of window length translation step. In each backtracking window, the spatial distribution dispersion of the location description is calculated to obtain the spatial dispersion pattern corresponding to the backtracking window. The spatial dispersion patterns of all backtracking windows are compared across windows to generate a spatial dispersion contraction trend indicator. In each backtracking window, the cumulative turning angle of the displacement segment formed by adjacent position descriptions is extracted to obtain the path detour description corresponding to that backtracking window. The path detour descriptions of all backtracking windows are compared across windows to generate an upward trend indicator of path detour. For each backtracking window, the ratio of the collection time span of all location descriptions within the window to the number of location descriptions appearing is calculated to obtain the location update density corresponding to the backtracking window. The location update density of all backtracking windows is compared across windows to generate an indicator of the trend of update density change. The spatial distribution shrinkage trend indicator, the path detourness increase trend indicator, and the update density change direction indicator are all input into the tendency fusion logic. The tendency fusion logic predefines the residence tendency level corresponding to different combinations of the three indicators. The dwell tendency level that matches the current combination of the three indications is output from the tendency fusion logic, and is used as a description of the dwell tendency of the target object in the current area.
6. The regional sound field adjustment method based on an inductive directional sound source according to claim 5, characterized in that, In each backtracking window, the cumulative turning angle of the displacement segment formed by adjacent position descriptions is extracted to obtain the path detour description corresponding to that backtracking window. A cross-window comparison is then performed on the path detour descriptions of all backtracking windows to generate a path detour upward trend indicator, including: Obtain the sequence of position descriptions arranged in chronological order within the current backtracking window, connect each position description with its immediately following position description to form a displacement segment, and record each displacement segment as a displacement segment vector containing direction angle and length; Calculate the directional deflection angle of each displacement segment vector relative to its immediate preceding displacement segment vector, map the directional deflection angle to a steering angle with a positive or negative sign, assign a positive sign to clockwise deflection and a negative sign to counterclockwise deflection, and generate a steering angle sign sequence. The steering angle symbol sequence is divided into consecutive segments with the same sign. The absolute cumulative amount of steering angle in each consecutive segment with the same sign and the time duration of the segment are calculated. The switching frequency of alternating segments with the same sign is extracted. The ratio of the absolute cumulative amount of each consecutive segment with the same number to the duration of that segment is taken as the turning density of that segment. Then, the turning density of all consecutive segments with the same number in the current backtracking window is statistically aggregated at the window level to obtain the representative value of the window turning density. At the same time, the switching frequency of alternating segments with the same number is taken as the window switching frequency. The representative value of the window turning density and the window switching frequency are normalized and then fused to generate the path detour description of the backtracking window. The path detour descriptions of adjacent backtracking windows are compared one by one along the time progression direction. When the path detour description shows a monotonically increasing trend and continues for at least a preset number of consecutive backtracking windows, it is determined that an upward trend in path detour has been detected. After determining that an upward trend in path detourness has been detected, the turning angle symbol sequence of each backtracking window within the increasing interval is further verified to see whether the switching frequency of the same-sign segment increases synchronously. When the switching frequency of the same-sign segment increases synchronously with the path detourness description, it is confirmed as the upward trend indication of path detourness and output.
7. The regional sound field adjustment method based on an inductive directional sound source according to claim 5, characterized in that, For each backtracking window, the ratio of the acquisition time span of all location descriptions within the window to the number of occurrences of each location description is calculated to obtain the location update density corresponding to that backtracking window. A cross-window comparison of the location update density of all backtracking windows is then performed to generate an indicator of the trend of update density changes, including: Within each backtracking window, the time interval between the acquisition times described by each adjacent location is extracted to form a location update time interval sequence. The interval distribution pattern of the location update time interval sequence is analyzed to distinguish between dense interval clusters and sparse interval clusters. The average time interval within the densely spaced cluster and the sparsely spaced cluster are calculated respectively. The average time interval of the densely spaced cluster and the average time interval of the sparsely spaced cluster are compared by span to generate the inter-cluster interval span ratio. The position update time interval sequence is divided into several continuous segments according to time sequence. The fluctuation degree of the time interval within each continuous segment is counted, and the transmission direction of the fluctuation degree between continuous segments is detected to generate the interval fluctuation transmission trend. By combining the inter-cluster interval span ratio and the interval fluctuation transmission trend, it is determined whether the position update in the current backtracking window shows an intermittent contraction trend. When the inter-cluster interval span ratio is higher than the preset inter-cluster convergence threshold and the interval fluctuation transmission trend points to the narrowing of fluctuation, a contraction determination indicator is generated. All backtracking windows are traversed one by one along the sliding direction. The shrinkage determination identifier of each backtracking window, together with its inter-cluster interval span ratio, constitutes the density state description of the backtracking window. The density state descriptions of adjacent backtracking windows are transferred and compared to detect the occurrence frequency and continuity of the shrinkage determination identifier. When the number of backtracking windows with consecutive contraction determination marks exceeds a preset threshold for the number of consecutive windows, and the density of position updates in each backtracking window within the consecutive segment shows a monotonically decreasing trend, an update density change trend indicator is generated and output. The update density change trend indicator indicates that the position updates are becoming denser and that this trend is continuous.
8. The method for regional sound field adjustment based on an inductive directional sound source according to claim 1, characterized in that, When the dwell tendency description reaches a preset tendency benchmark, and the output progress of the first audio content reaches any semantic breakpoint among the semantic segmentation markers, the end of the output sampling sequence of the first audio content is continuously spliced with the beginning of the output sampling sequence of the second audio content in the output buffer to form a seamless audio stream, including: When the dwell tendency description reaches the tendency benchmark, a dwell stability indication is generated, and the output progress of the first audio content is continuously monitored to see if it touches any semantic breakpoint in the semantic segment division marker. At the moment when the output progress reaches the semantic breakpoint, the sampling point of the preset connection preprocessing length is rolled back as the splicing preparation start point. Starting from the splicing preparation start point, the trailing sampling sequence of the first audio content and the leading sampling sequence of the second audio content are simultaneously acquired. From the trailing sample sequence of the first audio content and the leading sample sequence of the second audio content, the amplitude envelopes of their respective signals are stripped to obtain the trailing envelope contour and the leading envelope contour. The trailing envelope contour is extrapolated and extended to generate a naturally decaying extension segment of the trailing envelope; the leading envelope contour is predicted to start and generate an expected climbing segment of the leading envelope. In the overlapping area of the natural decay extension segment of the trailing envelope and the expected climbing segment of the leading envelope, a smooth mixing interval is defined with the intersection of the two envelope curves as the center, and amplitude cross-variation mixing is performed within the smooth mixing interval to obtain the envelope control curve of the mixing transition segment. The envelope control curve of the hybrid transition segment is applied to the corresponding sampling points of the first audio content trailing sampling sequence and the second audio content leading sampling sequence to synthesize a transition segment. The sampling sequence of the first audio content before the start of the transition preprocessing, the transition segment, and the sampling sequence of the second audio content after the transition interval are sequentially concatenated to obtain the uninterrupted connected audio stream. Specifically, in the overlapping region of the natural decay extension segment of the trailing envelope and the expected ascent segment of the leading envelope, a smooth mixing interval is defined centered on the intersection of the two envelope curves, and amplitude cross-variable mixing is performed within the smooth mixing interval to obtain the envelope control curve of the mixing transition segment, including: In the overlapping region of the natural decay extension of the trailing envelope and the expected rise of the leading envelope, the intersection point of the two envelope curves is searched along the time axis, and this intersection point is taken as the amplitude mixing center point. Based on the amplitude mixing center point, a preset mixing half-interval duration is extended forward and backward respectively, and the closed interval from the end point of the forward extension to the end point of the backward extension is taken as the smooth mixing interval. Within the smooth mixing interval, the envelope curve of the naturally decaying extension segment of the trailing envelope and the envelope curve of the expected climbing segment of the leading envelope are compared point by point at each time to identify the amplitude dominance relationship between the two curves at each time. At the beginning of the smooth mixing interval, the mixing weight is fully allocated to the naturally decaying extension segment of the trailing envelope; as time progresses towards the amplitude mixing center point, the mixing weight ratio of the expected climbing segment of the leading envelope is gradually increased; at the amplitude mixing center point, the mixing weights of the two envelopes are equal; after passing the amplitude mixing center point, the mixing weight ratio of the expected climbing segment of the leading envelope continues to increase until the mixing weight at the end of the smooth mixing interval is fully allocated to the expected climbing segment of the leading envelope; At each moment, the envelope amplitude values of the naturally decaying extension segment of the trailing envelope and the expected climbing segment of the leading envelope are weighted and mixed according to the aforementioned mixing weight ratio to generate the mixed envelope amplitude value at that moment. The mixing envelope amplitude values at all times within the smooth mixing interval are collected to form the envelope control curve of the mixing transition segment.
9. A server, characterized in that, include: processor; And a memory, wherein the memory stores computer-readable code that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.