keyboard_arrow_up
Energy-Efficient On-Chip Adaptive Spike Sorting And Multiresolution DWT Via Reconfigurable MAC And Convolution Accelerators In 22 NM FDSOI

Authors

Annika Weisse, Richard M. George, Stephan Henker and Christian Mayr, Technical University of Dresden (TUD), Germany

Abstract

This paper presents a real-timenovel, energy-efficient signal processing architecture for real-time on-chip processing of electrode recordings from neuronal tissue, with the overarching objective of enabling fully autonomous, adaptive neural recording systems. The primary goal is to overcome the limitations of traditional off-chip data transfer and post-processing by integrating high-performance, configurable hardware accelerators directly onto the same chip as the processing core. To achieve this, we introduce In combination with ana IBEX RISC-V based IBEX CPU augmented with , two dedicatedspecialized hardware accelerators are implemented for state-of-the-art on-chip neural signal processing tasks: a highly configurable Multiply-Accumulate (MAC) unit and a Convolution engine (CONV), both highly configurable and optimized for signed fixed-point arithmetic. The MAC unit significantly enhances on-chip processing efficiency by offloadingoffloads computationally intensive operations from the CPU - such as those in spike detection and feature extraction - thereby significantly reducing CPU load and energy consumption. To enable fully autonomous, on-chip spike sorting, a lightweight Principal Component Analysis (PCA) routine has been implemented directly on the CPU, allowing for on-chip feature extraction and continuous model retraining in response to dynamic changes such as . This architecture enables on-chip feature extraction and training. This on-chip training capability eliminates the need for offline preprocessing and enables real-time adaptation to changes in the recording due to electrode drift ,or artifacts, etc. This capability eliminates the need for offline preprocessing and supports closed-loop, adaptive neural interfaces. The convolution engine further enhances the system's versatility by enabling multiresolution time-frequency analysis through band-selective filtering and Discrete Wavelet Transform (DWT) decomposition. These capabilities collectively support advanced signal processing tasks - includingWe demonstrate the flexibility and power of the accelerators by implementing multiresolution spectral analysis and competitive spike sorting and spectral analysis - directly on the chip., Implemented in 22 nm fully-depleted silicon-on-insulator (FDSOI) technology, the System-on-Chip (SoC) operates at ultra-low voltage (0.55 V), leveraging the PCA-based feature extraction pipeline. Employing the MAC unit for inference reduces the energy consumption from 1.43 μJ/spike to 1.09 μJ/spike, representingachieving a 23.8 % improvement in energy efficiency compared to a pure CPU implementation(1.09 µJ/spike vs. 1.43 µJ/spike) compared to a pure CPU approach, and reducing . The Convolution engine supports band-selective filtering and multiresolution Discrete Wavelet Transform (DWT) decomposition, facilitating on-chip time-frequency analysis and enabling up to 75 % reduction in off-chip data bandwidth by up to 75 %. The chosen processing tasks underline the utility of convolution and filter accelerators in particular, when it comes to designing a flexible and hardware-efficient computing platform for neuronal signal analysis. The presented System-on-Chip implemented in 22 nm fully-depleted silicon-on-insulator (FDSOI) technology, facilitating ultra-low-voltage operation at 0.55 V. The results demonstrate that the integration of dedicated, flexible accelerators for convolution and MAC operations is essential for building scalable, low-power, ad intelligent neural interfaces capable of real-time, autonomous signal analysis.

Keywords

Neural Signal Processing, Neural Prosthetics, System-on-Chip, Spike Sorting, Multiply-Accumulate Unit, Finite Impulse Response Filter, Principal Component Analysis, Convolution, Multiresolution Discrete Wavelet Transform.

Full Text  Volume 16, Number 12