latencies to main memory by overlapping long-latency loads in stalled threads with useful computation